Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion .github/workflows/R-CMD-check.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -2,7 +2,7 @@
# Need help debugging build failures? Start at https://github.com/r-lib/actions#where-to-find-help
on:
push:
branches: [main, master]
branches: [main]
pull_request:
schedule:
- cron: '0 7 * * *'
Expand Down
10 changes: 6 additions & 4 deletions DESCRIPTION
Original file line number Diff line number Diff line change
@@ -1,7 +1,7 @@
Package: tongfen
Type: Package
Title: Make Data Based on Different Geographies Comparable
Version: 0.3.8
Version: 0.3.9
Authors@R: c(
person("Jens", "von Bergmann", email = "jens@mountainmath.ca", role = c("aut", "cre"), comment = "creator and maintainer"))
Description: Several functions to allow comparisons of data across different geographies, in particular for Canadian census data from different censuses.
Expand All @@ -11,17 +11,18 @@ ByteCompile: yes
LazyData: true
NeedsCompilation: no
Imports:
dplyr (>= 1.0),
dplyr (>= 1.1.0),
tidyr (>= 1.0),
sf,
tibble,
rlang,
purrr,
stringr,
readr,
nanoparquet,
tools,
utils,
lifecycle
RoxygenNote: 7.3.3
Suggests:
knitr,
rmarkdown,
Expand All @@ -34,11 +35,12 @@ Suggests:
readxl,
scales,
microbenchmark,
testthat (>= 3.0.0)
testthat (>= 3.2.0)
VignetteBuilder: knitr, rmarkdown
URL: https://github.com/mountainMath/tongfen, https://mountainmath.github.io/tongfen/
BugReports: https://github.com/mountainMath/tongfen/issues
Language: en-US
RdMacros: lifecycle
Depends:
R (>= 4.1)
Config/roxygen2/version: 8.1.0
4 changes: 4 additions & 0 deletions NAMESPACE
Original file line number Diff line number Diff line change
Expand Up @@ -14,9 +14,13 @@ export(meta_for_additive_variables)
export(meta_for_ca_census_vectors)
export(proportional_reaggregate)
export(tongfen_aggregate)
export(tongfen_anomaly_joins)
export(tongfen_ca_census_ct)
export(tongfen_detect_anomalies)
export(tongfen_estimate)
export(tongfen_estimate_ca_census)
export(tongfen_join_correspondence)
export(tongfen_join_regions)
export(tongfen_tag_largest_overlap)
import(dplyr)
import(rlang)
Expand Down
60 changes: 57 additions & 3 deletions NEWS.md
Original file line number Diff line number Diff line change
@@ -1,3 +1,57 @@
# tongfen v.0.3.9
## Major changes
- new experimental functions to detect and correct for likely geocoding anomalies in timelines on a
common geography, where the same dwellings got assigned to different neighbouring regions in
different years. This shows up as a surprising drop in one region that is offset by a jump in a
neighbouring region. `tongfen_detect_anomalies` lists the candidate regions,
`tongfen_anomaly_joins` determines the regions to join, `tongfen_join_regions` joins them in
data that already is on a common geography and `tongfen_join_correspondence` joins them in a
correspondence for use with `tongfen_aggregate`. See the new "Geocoding anomalies in TongFen
timelines" vignette for details
- StatCan correspondence files are now downloaded as parquet files from a mirror, Statistics Canada
put the original files behind a browser check that blocks programmatic downloads, which broke
`method = "statcan"`. Cached files are checked against the mirror once per session and downloaded
again if they changed. The mirror location can be changed via the `tongfen.statcan_correspondence_url`
option. Previously cached `statcan_correspondence_*.csv` files in the tongfen cache directory are
no longer used and can be removed
## Minor changes
- combining correspondences across three or more datasets no longer runs into a cross join when the
correspondences happen to be ordered so that consecutive ones don't share a geographic identifier.
This gave a dplyr deprecation warning and needlessly large intermediate tables, results are unchanged
- fix `proportional_reaggregate` giving wrong results when the finer level data already has values
for the categories to reaggregate. The values were compared to the parent total across all
categories instead of per category, and a missing value in one child region discarded the
existing values of all its siblings
- fix `tongfen_estimate` underestimating values when the intersection of a source and a target
region is a geometry collection, i.e. several polygons joined by a shared boundary line
- `tongfen_estimate` with `na.rm = FALSE` no longer returns `NA` for target regions that only
share a boundary with a source region with missing values
- `tongfen_estimate` now returns `NA` for target regions that don't overlap the source instead of
erroring out when none of them do, and gives a clear error when `target` already has a column
named like one of the variables to estimate
- fix `tongfen_aggregate` and `aggregate_data_with_meta` scaling averages by their parent variable
more than once when the metadata lists the same variable name for several datasets, as is common
with US census data. `tongfen_aggregate` now only uses the metadata of the dataset being
aggregated, conflicting aggregation rules for the same variable are an error
- averages aggregated with `na.rm = TRUE` are now taken over the regions that have a value. The
parent variable of regions with a missing average still counted toward the total the average
was divided by, pulling the result toward zero. This affects `aggregate_data_with_meta`,
`tongfen_aggregate`, `tongfen_estimate` and the functions built on them
- fix "Average to" variables like percentage changes overwriting each other's base when several of
them share a parent variable. In that case the base columns in the result are named after the
variable (`base_<variable>`) instead of the parent
- `estimate_tongfen_correspondence` no longer requires the geometry column to be named `geometry`
- `refresh = TRUE` in `get_tongfen_correspondence_ca_census` and `get_tongfen_ca_census` now also
refreshes the cached StatCan correspondence files
- US Census Bureau relationship files are downloaded to a temporary file first, an interrupted
download no longer leaves a broken file in the cache. They are now cached in the same place as
the StatCan correspondence files, also honouring the `tongfen.cache_path` environment variable
and the `custom_data_path` option
- `tongfen_estimate_ca_census` returns its result visibly
- documentation fixes, among others the `tongfen_aggregate` example now passes a named list of
datasets matching the metadata
- requires dplyr 1.1.0 or newer, which the package already relied on

# tongfen v.0.3.8
## Breaking changes
- `get_tongfen_ca_census` now honours its `base_geo`, `na.rm`, `tolerance`, `crs` and
Expand Down Expand Up @@ -68,12 +122,12 @@
- squish several edge case bugs

# tongfen v.0.3.6
## Major changs
## Major changes
- better downsampling that can also accommodate averages
- performance improvements
## Minor changes
- better documentation
- allow for datasets vartiables by census year for canadian data
- allow for datasets variables by census year for Canadian data
- fix issue where some metadata might get duplicated

# tongfen 0.3.2
Expand All @@ -85,7 +139,7 @@
## Major changes
- Added `tongfen_estimate_ca_census` function for new CensusMapper endpoint, tying into new {cancensus} functionality.
## Minor changes
- Custom impelementation of `tongfen_etimate` for finer control
- Custom implementation of `tongfen_estimate` for finer control
- Fix compatibility issue with changes in {sf} package

# tongfen 0.3
Expand Down
86 changes: 75 additions & 11 deletions R/helpers.R
Original file line number Diff line number Diff line change
Expand Up @@ -14,6 +14,57 @@ tongfen_cache_dir <- function(){
tempdir()
}

tongfen_session <- new.env(parent=emptyenv())

# ETag of a remote file, NULL if it can't be determined, e.g. when offline
remote_etag <- function(url){
headers <- tryCatch(suppressWarnings(curlGetHeaders(url)),error=function(e) NULL)
if (is.null(headers) || !identical(attr(headers,"status"),200L)) return(NULL)
etag <- grep("^etag:",headers,ignore.case=TRUE,value=TRUE)
if (length(etag)==0) return(NULL)
gsub("^etag:\\s*|\"|\\s+$","",etag[length(etag)],ignore.case=TRUE)
}

# location of the cached US Census Bureau relationship files
us_cache_dir <- function(cache_path=NULL){
file.path(nullify_blank(cache_path) %||% tongfen_cache_dir(),"us_data")
}

# Download a remote file to the local path unless the local copy is still current.
# The ETag of the downloaded file is kept next to the cached file and compared to the remote ETag
# the first time the file is requested in a session, the file is only downloaded again if it changed.
# Files that never change don't need to be checked against the remote, with `check_remote=FALSE`
# the cached file is used as is. The download goes to a temporary file first so that an
# interrupted download does not leave a broken file in the cache.
cached_download <- function(url,path,refresh=FALSE,check_remote=TRUE){
etag_path <- paste0(path,".etag")
cached <- file.exists(path) && !refresh
if (cached && (!check_remote || isTRUE(tongfen_session[[url]]))) return(path)
etag <- if (check_remote) remote_etag(url)
if (cached) {
if (is.null(etag)) {
message(paste0("Could not check ",url," for updates, using cached version."))
}
local_etag <- if (file.exists(etag_path)) readLines(etag_path,n=1,warn=FALSE)
if (is.null(etag) || identical(etag,local_etag)) {
tongfen_session[[url]] <- TRUE
return(path)
}
}
if (!dir.exists(dirname(path))) dir.create(dirname(path),recursive=TRUE)
tmp <- tempfile(tmpdir=dirname(path))
on.exit(unlink(tmp))
utils::download.file(url,tmp,mode="wb",quiet=TRUE)
# S3 ETags of files that were not uploaded in parts are the md5 checksum of the file
if (!is.null(etag) && grepl("^[0-9a-f]{32}$",etag) && !identical(unname(tools::md5sum(tmp)),etag)) {
stop(paste0("Download of ",url," is corrupted, please try again."))
}
file.copy(tmp,path,overwrite=TRUE)
if (is.null(etag)) unlink(etag_path) else writeLines(etag,etag_path)
tongfen_session[[url]] <- TRUE
path
}

inner_join_tongfen_correspondence <- function(data,correspondence,link){
data %>%
inner_join(correspondence %>%
Expand Down Expand Up @@ -140,6 +191,15 @@ assert <- function (expr, error) {
if (! expr) stop(error, call. = FALSE)
}

# Pairs of intersecting geometries as row indices into `x` and `y`. The sparse index
# list is turned into a tibble directly, `as.data.frame()` on an empty result drops
# the columns we join on.
intersects_pairs <- function(x, y) {
m <- sf::st_intersects(x, y, sparse = TRUE)
tibble(row.id = rep(seq_along(m), lengths(m)),
col.id = as.integer(unlist(m)))
}


# Dissolve the geometries of `data` by `grouping_var`, the geometric equivalent
# of `summarize()`. Groups holding a single geometry - the bulk of the groups
Expand Down Expand Up @@ -196,18 +256,22 @@ aggregate_correspondences <- function(correspondences){
select(!matches("Tongfen") | matches("TongfenMethod"))
}
# compute full correspondence, smallest table first to keep intermediate
# join results as small as possible
index_order <- correspondences %>% lapply(nrow) %>% unlist() %>% order()

correspondence <- correspondences[[index_order[1]]] %>%
clean_correspondence_names()
if (length(correspondences)>1) for (index in index_order[-1]) {
c <- correspondences[[index]] %>%
clean_correspondence_names()
match_columns <- intersect(names(correspondence),names(c))
match_columns <- match_columns[!grepl("TongfenMethod",match_columns)]
correspondence <- inner_join(correspondence,c,by=match_columns) %>%
# join results as small as possible, but only join tables that share an identifier
# with the tables joined so far, joining unrelated tables gives a cross join
remaining <- correspondences[order(vapply(correspondences,nrow,integer(1)))] %>%
lapply(clean_correspondence_names)
correspondence <- remaining[[1]]
remaining <- remaining[-1]
while (length(remaining)>0) {
match_columns <- lapply(remaining,function(c) {
match_columns <- intersect(names(correspondence),names(c))
match_columns[!grepl("TongfenMethod",match_columns)]
})
index <- which(lengths(match_columns)>0)[1]
if (is.na(index)) stop("Correspondences can't be combined, they don't share a common geographic identifier.")
correspondence <- inner_join(correspondence,remaining[[index]],by=match_columns[[index]]) %>%
unique()
remaining <- remaining[-index]
}

method_columns <- names(correspondence)[grepl("TongfenMethod",names(correspondence))]
Expand Down
2 changes: 1 addition & 1 deletion R/tonfen_deprecated.R
Original file line number Diff line number Diff line change
@@ -1,4 +1,4 @@
#' Check geographic integrety
#' Check geographic integrity
#'
#' @description
#' \lifecycle{deprecated}
Expand Down
Loading
Loading