diff --git a/.github/workflows/R-CMD-check.yaml b/.github/workflows/R-CMD-check.yaml
index bd2aadc..472ced2 100644
--- a/.github/workflows/R-CMD-check.yaml
+++ b/.github/workflows/R-CMD-check.yaml
@@ -2,7 +2,7 @@
# Need help debugging build failures? Start at https://github.com/r-lib/actions#where-to-find-help
on:
push:
- branches: [main, master]
+ branches: [main]
pull_request:
schedule:
- cron: '0 7 * * *'
diff --git a/DESCRIPTION b/DESCRIPTION
index c014060..2631df5 100644
--- a/DESCRIPTION
+++ b/DESCRIPTION
@@ -1,7 +1,7 @@
Package: tongfen
Type: Package
Title: Make Data Based on Different Geographies Comparable
-Version: 0.3.8
+Version: 0.3.9
Authors@R: c(
person("Jens", "von Bergmann", email = "jens@mountainmath.ca", role = c("aut", "cre"), comment = "creator and maintainer"))
Description: Several functions to allow comparisons of data across different geographies, in particular for Canadian census data from different censuses.
@@ -11,7 +11,7 @@ ByteCompile: yes
LazyData: true
NeedsCompilation: no
Imports:
- dplyr (>= 1.0),
+ dplyr (>= 1.1.0),
tidyr (>= 1.0),
sf,
tibble,
@@ -19,9 +19,10 @@ Imports:
purrr,
stringr,
readr,
+ nanoparquet,
+ tools,
utils,
lifecycle
-RoxygenNote: 7.3.3
Suggests:
knitr,
rmarkdown,
@@ -34,7 +35,7 @@ Suggests:
readxl,
scales,
microbenchmark,
- testthat (>= 3.0.0)
+ testthat (>= 3.2.0)
VignetteBuilder: knitr, rmarkdown
URL: https://github.com/mountainMath/tongfen, https://mountainmath.github.io/tongfen/
BugReports: https://github.com/mountainMath/tongfen/issues
@@ -42,3 +43,4 @@ Language: en-US
RdMacros: lifecycle
Depends:
R (>= 4.1)
+Config/roxygen2/version: 8.1.0
diff --git a/NAMESPACE b/NAMESPACE
index 1ba1cc6..16f2796 100644
--- a/NAMESPACE
+++ b/NAMESPACE
@@ -14,9 +14,13 @@ export(meta_for_additive_variables)
export(meta_for_ca_census_vectors)
export(proportional_reaggregate)
export(tongfen_aggregate)
+export(tongfen_anomaly_joins)
export(tongfen_ca_census_ct)
+export(tongfen_detect_anomalies)
export(tongfen_estimate)
export(tongfen_estimate_ca_census)
+export(tongfen_join_correspondence)
+export(tongfen_join_regions)
export(tongfen_tag_largest_overlap)
import(dplyr)
import(rlang)
diff --git a/NEWS.md b/NEWS.md
index e9babd4..2c97b3e 100644
--- a/NEWS.md
+++ b/NEWS.md
@@ -1,3 +1,57 @@
+# tongfen v.0.3.9
+## Major changes
+- new experimental functions to detect and correct for likely geocoding anomalies in timelines on a
+ common geography, where the same dwellings got assigned to different neighbouring regions in
+ different years. This shows up as a surprising drop in one region that is offset by a jump in a
+ neighbouring region. `tongfen_detect_anomalies` lists the candidate regions,
+ `tongfen_anomaly_joins` determines the regions to join, `tongfen_join_regions` joins them in
+ data that already is on a common geography and `tongfen_join_correspondence` joins them in a
+ correspondence for use with `tongfen_aggregate`. See the new "Geocoding anomalies in TongFen
+ timelines" vignette for details
+- StatCan correspondence files are now downloaded as parquet files from a mirror, Statistics Canada
+ put the original files behind a browser check that blocks programmatic downloads, which broke
+ `method = "statcan"`. Cached files are checked against the mirror once per session and downloaded
+ again if they changed. The mirror location can be changed via the `tongfen.statcan_correspondence_url`
+ option. Previously cached `statcan_correspondence_*.csv` files in the tongfen cache directory are
+ no longer used and can be removed
+## Minor changes
+- combining correspondences across three or more datasets no longer runs into a cross join when the
+ correspondences happen to be ordered so that consecutive ones don't share a geographic identifier.
+ This gave a dplyr deprecation warning and needlessly large intermediate tables, results are unchanged
+- fix `proportional_reaggregate` giving wrong results when the finer level data already has values
+ for the categories to reaggregate. The values were compared to the parent total across all
+ categories instead of per category, and a missing value in one child region discarded the
+ existing values of all its siblings
+- fix `tongfen_estimate` underestimating values when the intersection of a source and a target
+ region is a geometry collection, i.e. several polygons joined by a shared boundary line
+- `tongfen_estimate` with `na.rm = FALSE` no longer returns `NA` for target regions that only
+ share a boundary with a source region with missing values
+- `tongfen_estimate` now returns `NA` for target regions that don't overlap the source instead of
+ erroring out when none of them do, and gives a clear error when `target` already has a column
+ named like one of the variables to estimate
+- fix `tongfen_aggregate` and `aggregate_data_with_meta` scaling averages by their parent variable
+ more than once when the metadata lists the same variable name for several datasets, as is common
+ with US census data. `tongfen_aggregate` now only uses the metadata of the dataset being
+ aggregated, conflicting aggregation rules for the same variable are an error
+- averages aggregated with `na.rm = TRUE` are now taken over the regions that have a value. The
+ parent variable of regions with a missing average still counted toward the total the average
+ was divided by, pulling the result toward zero. This affects `aggregate_data_with_meta`,
+ `tongfen_aggregate`, `tongfen_estimate` and the functions built on them
+- fix "Average to" variables like percentage changes overwriting each other's base when several of
+ them share a parent variable. In that case the base columns in the result are named after the
+ variable (`base_`) instead of the parent
+- `estimate_tongfen_correspondence` no longer requires the geometry column to be named `geometry`
+- `refresh = TRUE` in `get_tongfen_correspondence_ca_census` and `get_tongfen_ca_census` now also
+ refreshes the cached StatCan correspondence files
+- US Census Bureau relationship files are downloaded to a temporary file first, an interrupted
+ download no longer leaves a broken file in the cache. They are now cached in the same place as
+ the StatCan correspondence files, also honouring the `tongfen.cache_path` environment variable
+ and the `custom_data_path` option
+- `tongfen_estimate_ca_census` returns its result visibly
+- documentation fixes, among others the `tongfen_aggregate` example now passes a named list of
+ datasets matching the metadata
+- requires dplyr 1.1.0 or newer, which the package already relied on
+
# tongfen v.0.3.8
## Breaking changes
- `get_tongfen_ca_census` now honours its `base_geo`, `na.rm`, `tolerance`, `crs` and
@@ -68,12 +122,12 @@
- squish several edge case bugs
# tongfen v.0.3.6
-## Major changs
+## Major changes
- better downsampling that can also accommodate averages
- performance improvements
## Minor changes
- better documentation
-- allow for datasets vartiables by census year for canadian data
+- allow for datasets variables by census year for Canadian data
- fix issue where some metadata might get duplicated
# tongfen 0.3.2
@@ -85,7 +139,7 @@
## Major changes
- Added `tongfen_estimate_ca_census` function for new CensusMapper endpoint, tying into new {cancensus} functionality.
## Minor changes
-- Custom impelementation of `tongfen_etimate` for finer control
+- Custom implementation of `tongfen_estimate` for finer control
- Fix compatibility issue with changes in {sf} package
# tongfen 0.3
diff --git a/R/helpers.R b/R/helpers.R
index eb527c3..2e15131 100644
--- a/R/helpers.R
+++ b/R/helpers.R
@@ -14,6 +14,57 @@ tongfen_cache_dir <- function(){
tempdir()
}
+tongfen_session <- new.env(parent=emptyenv())
+
+# ETag of a remote file, NULL if it can't be determined, e.g. when offline
+remote_etag <- function(url){
+ headers <- tryCatch(suppressWarnings(curlGetHeaders(url)),error=function(e) NULL)
+ if (is.null(headers) || !identical(attr(headers,"status"),200L)) return(NULL)
+ etag <- grep("^etag:",headers,ignore.case=TRUE,value=TRUE)
+ if (length(etag)==0) return(NULL)
+ gsub("^etag:\\s*|\"|\\s+$","",etag[length(etag)],ignore.case=TRUE)
+}
+
+# location of the cached US Census Bureau relationship files
+us_cache_dir <- function(cache_path=NULL){
+ file.path(nullify_blank(cache_path) %||% tongfen_cache_dir(),"us_data")
+}
+
+# Download a remote file to the local path unless the local copy is still current.
+# The ETag of the downloaded file is kept next to the cached file and compared to the remote ETag
+# the first time the file is requested in a session, the file is only downloaded again if it changed.
+# Files that never change don't need to be checked against the remote, with `check_remote=FALSE`
+# the cached file is used as is. The download goes to a temporary file first so that an
+# interrupted download does not leave a broken file in the cache.
+cached_download <- function(url,path,refresh=FALSE,check_remote=TRUE){
+ etag_path <- paste0(path,".etag")
+ cached <- file.exists(path) && !refresh
+ if (cached && (!check_remote || isTRUE(tongfen_session[[url]]))) return(path)
+ etag <- if (check_remote) remote_etag(url)
+ if (cached) {
+ if (is.null(etag)) {
+ message(paste0("Could not check ",url," for updates, using cached version."))
+ }
+ local_etag <- if (file.exists(etag_path)) readLines(etag_path,n=1,warn=FALSE)
+ if (is.null(etag) || identical(etag,local_etag)) {
+ tongfen_session[[url]] <- TRUE
+ return(path)
+ }
+ }
+ if (!dir.exists(dirname(path))) dir.create(dirname(path),recursive=TRUE)
+ tmp <- tempfile(tmpdir=dirname(path))
+ on.exit(unlink(tmp))
+ utils::download.file(url,tmp,mode="wb",quiet=TRUE)
+ # S3 ETags of files that were not uploaded in parts are the md5 checksum of the file
+ if (!is.null(etag) && grepl("^[0-9a-f]{32}$",etag) && !identical(unname(tools::md5sum(tmp)),etag)) {
+ stop(paste0("Download of ",url," is corrupted, please try again."))
+ }
+ file.copy(tmp,path,overwrite=TRUE)
+ if (is.null(etag)) unlink(etag_path) else writeLines(etag,etag_path)
+ tongfen_session[[url]] <- TRUE
+ path
+}
+
inner_join_tongfen_correspondence <- function(data,correspondence,link){
data %>%
inner_join(correspondence %>%
@@ -140,6 +191,15 @@ assert <- function (expr, error) {
if (! expr) stop(error, call. = FALSE)
}
+# Pairs of intersecting geometries as row indices into `x` and `y`. The sparse index
+# list is turned into a tibble directly, `as.data.frame()` on an empty result drops
+# the columns we join on.
+intersects_pairs <- function(x, y) {
+ m <- sf::st_intersects(x, y, sparse = TRUE)
+ tibble(row.id = rep(seq_along(m), lengths(m)),
+ col.id = as.integer(unlist(m)))
+}
+
# Dissolve the geometries of `data` by `grouping_var`, the geometric equivalent
# of `summarize()`. Groups holding a single geometry - the bulk of the groups
@@ -196,18 +256,22 @@ aggregate_correspondences <- function(correspondences){
select(!matches("Tongfen") | matches("TongfenMethod"))
}
# compute full correspondence, smallest table first to keep intermediate
- # join results as small as possible
- index_order <- correspondences %>% lapply(nrow) %>% unlist() %>% order()
-
- correspondence <- correspondences[[index_order[1]]] %>%
- clean_correspondence_names()
- if (length(correspondences)>1) for (index in index_order[-1]) {
- c <- correspondences[[index]] %>%
- clean_correspondence_names()
- match_columns <- intersect(names(correspondence),names(c))
- match_columns <- match_columns[!grepl("TongfenMethod",match_columns)]
- correspondence <- inner_join(correspondence,c,by=match_columns) %>%
+ # join results as small as possible, but only join tables that share an identifier
+ # with the tables joined so far, joining unrelated tables gives a cross join
+ remaining <- correspondences[order(vapply(correspondences,nrow,integer(1)))] %>%
+ lapply(clean_correspondence_names)
+ correspondence <- remaining[[1]]
+ remaining <- remaining[-1]
+ while (length(remaining)>0) {
+ match_columns <- lapply(remaining,function(c) {
+ match_columns <- intersect(names(correspondence),names(c))
+ match_columns[!grepl("TongfenMethod",match_columns)]
+ })
+ index <- which(lengths(match_columns)>0)[1]
+ if (is.na(index)) stop("Correspondences can't be combined, they don't share a common geographic identifier.")
+ correspondence <- inner_join(correspondence,remaining[[index]],by=match_columns[[index]]) %>%
unique()
+ remaining <- remaining[-index]
}
method_columns <- names(correspondence)[grepl("TongfenMethod",names(correspondence))]
diff --git a/R/tonfen_deprecated.R b/R/tonfen_deprecated.R
index 9efe73c..372ed12 100644
--- a/R/tonfen_deprecated.R
+++ b/R/tonfen_deprecated.R
@@ -1,4 +1,4 @@
-#' Check geographic integrety
+#' Check geographic integrity
#'
#' @description
#' \lifecycle{deprecated}
diff --git a/R/tongfen.R b/R/tongfen.R
index 5640454..f81c484 100644
--- a/R/tongfen.R
+++ b/R/tongfen.R
@@ -5,8 +5,8 @@
#'
#' Generates metadata to be used in tongfen_aggregate. Variables need to be additive like counts.
#'
-#' @param dataset identifier for the dataset contianing the variable
-#' @param variables (named) vecotor with additive variables
+#' @param dataset identifier for the dataset containing the variable
+#' @param variables (named) vector with additive variables
#' @return a tibble to be used in tongfen_aggregate
#' @export
#'
@@ -29,6 +29,21 @@ meta_for_additive_variables <- function(dataset,variables){
+# Averages are aggregated by scaling them by their parent variable, summing up, and dividing
+# by the summed up parent again. The parent of regions where the average is missing must not
+# count toward the total the sum gets divided by, so each average carries its own weight column.
+average_weight_name <- function(variables) {
+ if (length(variables)==0) return(character(0))
+ paste0("...weight_",variables)
+}
+
+average_weight_exprs <- function(to_scale,parent_lookup) {
+ lapply(setNames(to_scale, average_weight_name(to_scale)), \(col) {
+ parent <- as.name(unname(parent_lookup[col]))
+ rlang::expr(replace(as.numeric(!!parent),is.na(!!as.name(col)),NA_real_))
+ })
+}
+
cut_meta <- function(data,meta){
meta <- meta %>%
filter(.data$variable %in% names(data)|.data$label %in% names(data)) %>%
@@ -52,14 +67,16 @@ pre_scale <- function(data,meta,meta_var="data_var",quiet=FALSE) {
if (length(to_scale) > 0) {
+ weight_exprs <- average_weight_exprs(to_scale,parent_lookup)
scale_exprs <- lapply(setNames(to_scale, to_scale), \(col)
rlang::expr(!!as.name(col) * !!as.name(unname(parent_lookup[col]))))
- data <- data %>% mutate(!!!scale_exprs)
+ data <- data %>% mutate(!!!weight_exprs) %>% mutate(!!!scale_exprs)
}
data
}
+# divides by the weights added in `pre_scale` if they are still around, and by the parent otherwise
post_scale <- function(data,meta,meta_var="data_var") {
meta_name_lookup <- setNames(meta %>% pull(meta_var),meta$variable)
meta$parent_name <- meta_name_lookup[meta$parent]
@@ -67,9 +84,14 @@ post_scale <- function(data,meta,meta_var="data_var") {
to_scale <- filter(meta,.data$rule %in% c("Median","Average")) %>% pull(meta_var)
if (length(to_scale) > 0) {
- scale_exprs <- lapply(setNames(to_scale, to_scale), \(col)
- rlang::expr(!!as.name(col) / !!as.name(unname(parent_lookup[col]))))
- data <- data %>% mutate(!!!scale_exprs)
+ scale_exprs <- lapply(setNames(to_scale, to_scale), \(col) {
+ weight <- average_weight_name(col)
+ if (!(weight %in% names(data))) weight <- unname(parent_lookup[col])
+ rlang::expr(!!as.name(col) / !!as.name(weight))
+ })
+ data <- data %>%
+ mutate(!!!scale_exprs) %>%
+ select(-any_of(average_weight_name(to_scale)))
}
data
@@ -102,6 +124,17 @@ post_scale <- function(data,meta,meta_var="data_var") {
#'}
aggregate_data_with_meta <- function(data,meta,geo=FALSE,na.rm=TRUE,quiet=FALSE){
meta <- meta %>% filter(.data$variable %in% names(data))
+ # the same variable can be listed several times if meta spans several datasets,
+ # each variable must only be aggregated (and scaled) once
+ duplicates <- duplicated(meta$variable)
+ if (any(duplicates)) {
+ rules <- meta %>% select(any_of(c("variable","rule","parent","units"))) %>% unique()
+ ambiguous <- unique(rules$variable[duplicated(rules$variable)])
+ if (length(ambiguous)>0)
+ stop(paste0("Conflicting aggregation rules in metadata for ",paste0(ambiguous,collapse = ", "),
+ ", please only pass the metadata for the dataset that is being aggregated."))
+ meta <- meta[!duplicates,]
+ }
grouping_var=groups(data) %>% as.character
parent_lookup <- setNames(meta$parent,meta$variable)
to_scale <- filter(meta,.data$rule %in% c("Median","Average"))$variable
@@ -116,16 +149,25 @@ aggregate_data_with_meta <- function(data,meta,geo=FALSE,na.rm=TRUE,quiet=FALSE)
message(paste0("Can't TongFen medians, will approximate by treating as averages: ",paste0(median_vars,collapse = ", ")))
}
+ weight_variables <- average_weight_name(to_scale)
if (length(to_scale) > 0) {
+ weight_exprs <- average_weight_exprs(to_scale,parent_lookup)
scale_exprs <- lapply(setNames(to_scale, to_scale), \(col)
rlang::expr(!!as.name(col) * !!as.name(unname(parent_lookup[col]))))
- data <- data %>% mutate(!!!scale_exprs)
+ data <- data %>% mutate(!!!weight_exprs) %>% mutate(!!!scale_exprs)
}
+ # the base an "Average to" variable gets scaled by depends on the variable, variables
+ # sharing a parent each need their own base column
+ scale_from_parents <- unname(parent_lookup[to_scale_from])
+ shared_parent <- scale_from_parents %in% scale_from_parents[duplicated(scale_from_parents)]
+ base_vectors <- setNames(paste0("base_",ifelse(shared_parent,to_scale_from,scale_from_parents)),
+ to_scale_from)
+
base_variables <- c()
for (x in to_scale_from) {
scale_type <- meta %>% filter(.data$variable==x) %>% pull(units) %>% as.character()
- base_vector <- paste0("base_",parent_lookup[x])
+ base_vector <- unname(base_vectors[x])
base_variables <- c(base_variables,base_vector)
if (scale_type=="Percentage ratio (0.0-1.0)") {
data <- data %>% mutate(!!base_vector:=!!as.name(parent_lookup[x])/(!!as.name(x)+1))
@@ -144,22 +186,24 @@ aggregate_data_with_meta <- function(data,meta,geo=FALSE,na.rm=TRUE,quiet=FALSE)
data <- left_join(summarize_geometry_by_group(data,grouping_var),
data %>%
sf::st_set_geometry(NULL) %>%
- summarize_at(meta$variable,sum,na.rm=na.rm),
+ summarize_at(c(meta$variable,weight_variables),sum,na.rm=na.rm),
by=grouping_var)
} else {
- data <- data %>% summarize_at(meta$variable,sum,na.rm=na.rm)
+ data <- data %>% summarize_at(c(meta$variable,weight_variables),sum,na.rm=na.rm)
}
if (length(to_scale) > 0) {
scale_exprs <- lapply(setNames(to_scale, to_scale), \(col)
- rlang::expr(!!as.name(col) / !!as.name(unname(parent_lookup[col]))))
- data <- data %>% mutate(!!!scale_exprs)
+ rlang::expr(!!as.name(col) / !!as.name(average_weight_name(col))))
+ data <- data %>%
+ mutate(!!!scale_exprs) %>%
+ select(-all_of(weight_variables))
}
# Optimized: Vectorized division by base vectors
if (length(to_scale_from) > 0) {
for (x in to_scale_from) {
- base_vector <- paste0("base_", parent_lookup[x])
+ base_vector <- unname(base_vectors[x])
data[[x]] <- data[[x]] / data[[base_vector]]
}
}
@@ -186,9 +230,12 @@ rename_with_meta <- function(data,meta,ds=NULL){
#' @description
#' \lifecycle{maturing}
#'
-#' Aggregate variables secified in meta for several datasets according to correspondence.
+#' Aggregate variables specified in meta for several datasets according to correspondence.
#'
-#' @param data list of datasets to be aggregated
+#' @param data named list of datasets to be aggregated. The names identify the datasets, they are
+#' matched against the `geo_dataset` column in `meta` to pick the aggregation rules and labels
+#' for each dataset. Without names, or with names not found in `meta`, the rules for all
+#' datasets are applied and the variables keep their original names
#' @param correspondence correspondence data for gluing up the datasets
#' @param meta metadata containing aggregation rules as for example returned by `meta_for_ca_census_vectors`
#' @param base_geo identifier for which data element to base the final geography on,
@@ -200,17 +247,17 @@ rename_with_meta <- function(data,meta,ds=NULL){
#' @export
#'
#' @examples
-#' # aggregate census tract level 2006 population data on common gepgraphy build through
+#' # aggregate census tract level 2006 and 2016 population data on common geography built through
#' # correspondence from 2006 and 2016 census tracts in the City of Vancouver.
#' \dontrun{
#' regions <- list(CSD="5915022")
#' geo1 <- cancensus::get_census("CA06",regions=regions,geo_format='sf',level='CT')
#' geo2 <- cancensus::get_census("CA16",regions=regions,geo_format='sf',level='CT')
-#' meta <- meta_for_additive_variables("CA06","Population")
+#' meta <- meta_for_additive_variables(c("CA06","CA16"),"Population")
#' correspondence <- get_tongfen_correspondence_ca_census(geo_datasets=c('CA06','CA16'),
#' regions=regions,level='CT')
-#' result <- tongfen_aggregate(list(geo1 %>% rename(GeoUIDCA06=GeoUID),
-#' geo2 %>% rename(GeoUIDCA16=GeoUID)),correspondence,meta)
+#' result <- tongfen_aggregate(list(CA06=geo1 %>% rename(GeoUIDCA06=GeoUID),
+#' CA16=geo2 %>% rename(GeoUIDCA16=GeoUID)),correspondence,meta)
#'}
tongfen_aggregate <- function(data,correspondence,meta=NULL, base_geo = NULL, na.rm = TRUE){
data <- ensure_names(data)
@@ -239,7 +286,13 @@ tongfen_aggregate <- function(data,correspondence,meta=NULL, base_geo = NULL, na
by=match_column) %>%
group_by(.data$TongfenID,.data$TongfenUID)
if (!is.null(meta)) {
- d <- d %>% aggregate_data_with_meta(meta,na.rm=na.rm)
+ # only use the aggregation rules for this dataset if meta distinguishes datasets,
+ # the same variable name can come with different rules in different datasets
+ ds_meta <- meta
+ if (ds %in% as.character(meta$geo_dataset)) {
+ ds_meta <- meta %>% filter(as.character(.data$geo_dataset)==ds)
+ }
+ d <- d %>% aggregate_data_with_meta(ds_meta,na.rm=na.rm)
} else {
if ("sf" %in% class(d)) {
d <- summarize_geometry_by_group(d,c("TongfenID","TongfenUID"))
@@ -291,7 +344,7 @@ tongfen_aggregate <- function(data,correspondence,meta=NULL, base_geo = NULL, na
#' @param geo_match A named string informing on what column names to match data and parent_data
#' @param categories Vector of column names to re-aggregate
#' @param base Column name to use for proportional weighting when re-aggregating, or named vector with column name for each category.
-#' Categries that should be re-aggregated as means should be set to NA and will only be reaggregated if the base data has NA values.
+#' Categories that should be re-aggregated as means should be set to NA and will only be reaggregated if the base data has NA values.
#' @return dataframe with downsampled variables from parent_data
#' @keywords reaggregate proportionally wrt base variable
#' @export
@@ -397,7 +450,9 @@ proportional_reaggregate <- function(data,parent_data,geo_match,categories,base=
values_to="p_value")
if (vt %in% c("numeric","integer","integer64")) {
d_combined <- full_join(d_base,d_parent,by=c(geo_match,"category"="category")) %>%
- mutate(s_value=sum(.data$value),.by=names(geo_match)) %>%
+ # the parent value gets compared to what the children of the same category
+ # already add up to, missing child values count as zero
+ mutate(s_value=sum(.data$value,na.rm=TRUE),.by=c(names(geo_match),"category")) %>%
mutate(across(any_of(c("p_value","s_value")),\(x)coalesce(x,0))) %>%
mutate(value=case_when(.data$agg_type=="additive" ~ coalesce(.data$value,0) + .data$weight*(.data$p_value-.data$s_value),
is.na(.data$value) ~ .data$p_value,
@@ -422,7 +477,7 @@ proportional_reaggregate <- function(data,parent_data,geo_match,categories,base=
select(-any_of(id))
}
-#' Generate togfen correspondence for two geographies
+#' Generate tongfen correspondence for two geographies
#'
#' @description
#' \lifecycle{maturing}
@@ -465,42 +520,35 @@ estimate_tongfen_single_correspondence <- function(geo1,geo2,geo1_uid,geo2_uid,
if (robust) {
- if (!st_is_valid(geo1)) geo1 <- geo1 %>% st_make_valid()
- if (!st_is_valid(geo2)) geo2 <- geo2 %>% st_make_valid()
+ if (!isTRUE(all(st_is_valid(geo1)))) geo1 <- geo1 %>% st_make_valid()
+ if (!isTRUE(all(st_is_valid(geo2)))) geo2 <- geo2 %>% st_make_valid()
}
- robust_tolerance_buffer <- function(geo,geo_uid,tolerance,max_tries=20) {
+ # works on the geometries directly, the geometry column can go by any name
+ robust_tolerance_buffer <- function(geo,tolerance,max_tries=20) {
t <- tolerance
- d <- geo
- d$geometry=st_buffer(geo$geometry,-t)
+ geometry <- st_geometry(geo)
+ buffered <- st_buffer(geometry,-t)
count=0
- empties <- st_is_empty(d)
+ empties <- st_is_empty(buffered)
while (sum(empties) > 0 & count 0) {
stop("Unable to match within given tolerance, some geographies are too fine.")
}
- d
+ st_set_geometry(geo,buffered)
}
- cgeo1 <- geo1 %>% robust_tolerance_buffer(geo_uid = geo1_uid,tolerance = tolerance)
- cgeo2 <- geo2 %>% robust_tolerance_buffer(geo_uid = geo2_uid,tolerance = tolerance)
-
- # Both intersections are necessary (buffered cgeo1 vs geo2, and cgeo2 vs geo1). The
- # sparse index list is turned into a tibble directly, `as.data.frame()` on an empty
- # result drops the columns we join on.
- intersects_pairs <- function(x, y) {
- m <- st_intersects(x, y, sparse = TRUE)
- tibble(row.id = rep(seq_along(m), lengths(m)),
- col.id = as.integer(unlist(m)))
- }
+ cgeo1 <- geo1 %>% robust_tolerance_buffer(tolerance = tolerance)
+ cgeo2 <- geo2 %>% robust_tolerance_buffer(tolerance = tolerance)
- i1 <- intersects_pairs(cgeo1, geo2) %>%
+ # Both intersections are necessary (buffered cgeo1 vs geo2, and cgeo2 vs geo1).
+ i1 <-intersects_pairs(cgeo1, geo2) %>%
left_join(id1, by = c("row.id" = "id1")) %>%
left_join(id2, by = c("col.id" = "id2")) %>%
select(-"row.id",-"col.id")
@@ -518,7 +566,7 @@ estimate_tongfen_single_correspondence <- function(geo1,geo2,geo1_uid,geo2_uid,
correspondence
}
-#' Generate togfen correspondence for list of geographies
+#' Generate tongfen correspondence for list of geographies
#'
#' @description
#' \lifecycle{maturing}
@@ -614,7 +662,7 @@ estimate_tongfen_correspondence <- function(data,
-#' Check geographic integrety
+#' Check geographic integrity
#'
#' @description
#' \lifecycle{maturing}
@@ -626,7 +674,7 @@ estimate_tongfen_correspondence <- function(data,
#' simplified independently and differ in how water features are cut out, so a sizable area
#' mismatch does not by itself mean the regions were matched up incorrectly.
#'
-#' @param data alist of geogrpahic data of class sf
+#' @param data a list of geographic data of class sf
#' @param correspondence Correspondence table with columns the unique geographic identifiers for each of the
#' geographies and the TongfenID (and optionally TongfenUID and TongfenMethod)
#' returned by `estimate_tongfen_correspondence`.
diff --git a/R/tongfen_anomalies.R b/R/tongfen_anomalies.R
new file mode 100644
index 0000000..e4fcd48
--- /dev/null
+++ b/R/tongfen_anomalies.R
@@ -0,0 +1,563 @@
+# Geocoding of census data varies over time, the same dwelling units can get assigned to
+# different, neighbouring, regions in different years. In a timeline on a common geography
+# this shows up as a surprising drop in one region that is offset by a jump in a neighbouring
+# region. The functions in this file detect such patterns and join the affected regions,
+# extending the idea behind TongFen from changing boundaries to data that got assigned across
+# boundaries. See https://doodles.mountainmath.ca/posts/2024-07-26-geocoding-errors-in-aggregate-data/
+# for background, the implementation follows the method described there.
+
+anomaly_params <- function(rel_scale,abs_scale,p,surprise_cutoff,total_surprise_cutoff,
+ cutoff_fact,surprise_reduction_const,sum_fact) {
+ params <- list(rel_scale=rel_scale,abs_scale=abs_scale,p=p,
+ surprise_cutoff=surprise_cutoff,total_surprise_cutoff=total_surprise_cutoff,
+ cutoff_fact=cutoff_fact,surprise_reduction_const=surprise_reduction_const,
+ sum_fact=sum_fact)
+ valid <- vapply(params,\(x) is.numeric(x) && length(x)==1 && !is.na(x),logical(1))
+ assert(all(valid),paste0("Need a single number for ",paste0(names(params)[!valid],collapse=", "),"."))
+ assert(rel_scale>0 && abs_scale>0 && p>0,"rel_scale, abs_scale and p have to be positive.")
+ params
+}
+
+# Surprise of a change in a count, on a scale from 0 to 1. Only decreases are surprising, a
+# relative decrease of `rel_scale` and an absolute decrease of `abs_scale` each are half way
+# to full surprise, and it takes both for a change to be surprising. The relative change is
+# taken with respect to `base`, which does not have to be the count the change started from.
+anomaly_surprise <- function(change,base,rel_scale,abs_scale) {
+ decrease_surprise <- function(x,scale) {
+ r <- x
+ r[] <- 0
+ decrease <- is.finite(x) & x<=0
+ r[decrease] <- 1-0.5^(-x[decrease]/scale)
+ r
+ }
+ decrease_surprise(change/base,rel_scale)*decrease_surprise(change,abs_scale)
+}
+
+# One round of looking for regions to join. `V` is a matrix of counts with a row per region
+# and a column per year, `edges` a two column matrix with the row indices of neighbouring
+# regions. Returns the candidate regions, most surprising first, with the neighbour that
+# takes away most of the surprise when joined and if that is enough to join them.
+anomaly_round <- function(V,edges,params) {
+ nt <- ncol(V)
+ change <- V[,-1,drop=FALSE]-V[,-nt,drop=FALSE]
+ base <- V[,-nt,drop=FALSE]
+ count_surprise <- \(S) rowSums(S>params$surprise_cutoff)
+ total_surprise <- \(S) rowSums(S^params$p)^(1/params$p)
+
+ S <- anomaly_surprise(change,base,params$rel_scale,params$abs_scale)
+ count <- count_surprise(S)
+ total <- total_surprise(S)
+
+ index <- which(count>0 & total>params$total_surprise_cutoff)
+ index <- index[order(-count[index],-total[index],-index)]
+ result <- tibble(index=index,
+ surprise_count=as.integer(count[index]),
+ surprise_total=total[index],
+ period=max.col(S[index,,drop=FALSE],ties.method="first"),
+ neighbour=NA_integer_,
+ surprise_total_joined=NA_real_,
+ join=FALSE)
+
+ from <- c(edges[,1],edges[,2])
+ to <- c(edges[,2],edges[,1])
+ keep <- from %in% index
+ from <- from[keep]
+ to <- to[keep]
+ if (length(from)==0) return(result)
+
+ # Joining a region with flat counts lowers the surprise just by growing the denominator
+ # of the relative change. To keep the surprise comparable the change of the joined region
+ # is measured against the counts of the candidate region alone.
+ # A neighbour with unknown change can't make up for a surprising change.
+ change_neighbour <- change[to,,drop=FALSE]
+ change_neighbour[is.na(change_neighbour)] <- 0
+ S_joined <- anomaly_surprise(change[from,,drop=FALSE]+change_neighbour,
+ base[from,,drop=FALSE],params$rel_scale,params$abs_scale)
+ count_joined <- count_surprise(S_joined)
+ total_joined <- total_surprise(S_joined)
+
+ o <- order(from,count_joined,total_joined,to)
+ best <- o[!duplicated(from[o])]
+ best <- best[match(result$index,from[best])]
+
+ result$neighbour <- to[best]
+ result$surprise_total_joined <- total_joined[best]
+ join <- result$surprise_total_joined < params$cutoff_fact*result$surprise_total |
+ result$surprise_total-result$surprise_total_joined > params$surprise_reduction_const |
+ result$surprise_total_joined < params$cutoff_fact*params$sum_fact*
+ (result$surprise_total+total[result$neighbour])
+ result$join <- !is.na(join) & join
+ result
+}
+
+# Joins regions until there are no more pairs of regions left to join. Returns the group
+# each region ended up in and the round in which it first got joined to another region.
+anomaly_joins <- function(V,edges,params) {
+ membership <- seq_len(nrow(V))
+ round_joined <- rep(NA_integer_,nrow(V))
+ round <- 0L
+ repeat {
+ candidates <- anomaly_round(V,edges,params)
+ candidates <- candidates[candidates$join,]
+ # regions can only be part of one join per round, the most surprising ones go first
+ matched <- logical(nrow(V))
+ target <- seq_len(nrow(V))
+ for (r in seq_len(nrow(candidates))) {
+ i <- candidates$index[r]
+ j <- candidates$neighbour[r]
+ if (matched[i] || matched[j]) next
+ matched[c(i,j)] <- TRUE
+ target[max(i,j)] <- min(i,j)
+ }
+ if (!any(matched)) break
+
+ round <- round+1L
+ round_joined[matched[membership] & is.na(round_joined)] <- round
+
+ # contract the joined regions, the neighbours of a joined region are the neighbours
+ # of the regions it is made up of
+ target <- match(target,sort(unique(target)))
+ membership <- target[membership]
+ V <- rowsum(V,target)
+ edges <- cbind(target[edges[,1]],target[edges[,2]])
+ edges <- edges[edges[,1]!=edges[,2],,drop=FALSE]
+ edges <- unique(cbind(pmin(edges[,1],edges[,2]),pmax(edges[,1],edges[,2])))
+ }
+ list(membership=membership,round=round_joined)
+}
+
+# Neighbouring regions as two column matrix of row indices into `ids`, each pair of
+# neighbours is listed once.
+anomaly_neighbours <- function(data,ids,neighbours) {
+ if (is.null(neighbours)) {
+ assert("sf" %in% class(data),
+ "Need data of class sf to determine neighbouring regions, alternatively specify the neighbours.")
+ geometry <- sf::st_geometry(data)
+ pairs <- intersects_pairs(geometry,geometry)
+ from <- pairs$row.id
+ to <- pairs$col.id
+ } else if (is.data.frame(neighbours)) {
+ assert(ncol(neighbours)>=2,
+ "Need neighbours to have two columns with the identifiers of neighbouring regions.")
+ from <- match(as.character(neighbours[[1]]),ids)
+ to <- match(as.character(neighbours[[2]]),ids)
+ } else if (is.list(neighbours)) {
+ # neighbours list like the ones from the spdep package, regions without neighbours hold a 0
+ region_ids <- attr(neighbours,"region.id") %||% names(neighbours) %||% ids
+ assert(length(region_ids)==length(neighbours),
+ "Need neighbours to have an entry for each region.")
+ from <- rep(seq_along(neighbours),lengths(neighbours))
+ to <- as.integer(unlist(neighbours))
+ keep <- !is.na(to) & to>0
+ from <- match(as.character(region_ids)[from[keep]],ids)
+ to <- match(as.character(region_ids)[to[keep]],ids)
+ } else {
+ stop("Don't know how to interpret neighbours, need a table with pairs of identifiers or a neighbours list.")
+ }
+ keep <- !is.na(from) & !is.na(to) & from!=to
+ unique(cbind(pmin(from[keep],to[keep]),pmax(from[keep],to[keep])))
+}
+
+anomaly_input <- function(data,variables,id,neighbours) {
+ assert(is.character(variables) && length(variables)>=2,
+ "Need at least two variables making up the timeline to look for anomalies.")
+ assert(is.character(id) && length(id)==1,"Need id to be the name of the identifier column.")
+ missing_columns <- setdiff(c(id,variables),names(data))
+ assert(length(missing_columns)==0,
+ paste0("Did not find ",paste0(missing_columns,collapse=", ")," in data."))
+ d <- data %>% ungroup()
+ if ("sf" %in% class(d)) d <- d %>% sf::st_drop_geometry()
+ not_numeric <- variables[!vapply(variables,\(v) is.numeric(d[[v]]),logical(1))]
+ assert(length(not_numeric)==0,
+ paste0("Variables have to be numeric, got ",paste0(not_numeric,collapse=", "),"."))
+ id_values <- d[[id]]
+ ids <- as.character(id_values)
+ assert(!anyNA(ids) && !anyDuplicated(ids),
+ paste0("Need ",id," to uniquely identify the regions in data."))
+
+ V <- matrix(as.numeric(unlist(d[variables],use.names=FALSE)),ncol=length(variables))
+ edges <- anomaly_neighbours(data,ids,neighbours)
+
+ # Ties are broken by the order of the regions, sort by identifier so that the result
+ # does not depend on the order the regions come in.
+ o <- order(ids,method="radix")
+ position <- order(o)
+ edges <- cbind(position[edges[,1]],position[edges[,2]])
+ list(V=V[o,,drop=FALSE],
+ edges=cbind(pmin(edges[,1],edges[,2]),pmax(edges[,1],edges[,2])),
+ ids=ids[o],
+ id_values=id_values[o])
+}
+
+
+#' Detect likely geocoding anomalies in timelines on a common geography
+#'
+#' @description
+#' \lifecycle{experimental}
+#'
+#' TongFen is only as good as the geocoding that assigned the underlying data to geographic regions
+#' in the first place. Geocoding varies over time, and the same dwelling units, and the people living
+#' in them, can get assigned to different neighbouring regions in different years. In a timeline
+#' on a common geography this shows up as a surprising drop in one region that is offset by
+#' a corresponding jump in a neighbouring region.
+#'
+#' This function lists the candidate regions with surprising drops in the given count variable,
+#' together with the neighbouring region that takes away most of the surprise when both are joined.
+#' Use it to check for possible problems and to calibrate the parameters before joining regions with
+#' `tongfen_anomaly_joins`. Not all surprising drops are due to geocoding problems, a drop that is
+#' not complemented by a neighbouring region is likely real.
+#'
+#' The surprise of a change between two consecutive years ranges from 0 to 1. Only decreases are surprising,
+#' the surprise is the product of the surprise of the relative and of the absolute decrease, so that it takes
+#' a decrease that is large in both relative and absolute terms to be surprising. The total surprise of a
+#' region is the `p`-norm of the surprises across all changes in the timeline.
+#'
+#' A candidate region and its neighbour are flagged for joining if joining reduces the total surprise of the
+#' candidate region to below `cutoff_fact` times its total surprise, or reduces it by more than
+#' `surprise_reduction_const`, or reduces it to below `cutoff_fact * sum_fact` times the sum of
+#' the total surprises of both regions. To keep this comparable the surprise after joining is computed
+#' from the change of the joined regions relative to the counts of the candidate region alone.
+#'
+#' @param data data on a common geography, with one row per region, for example as returned by
+#' `tongfen_aggregate` or `get_tongfen_ca_census`. Needs to be of class sf unless `neighbours` is specified.
+#' @param variables names of the columns holding the timeline of a count variable like population or
+#' dwellings, in temporal order. Changes from or to a missing value are not surprising and don't make up for
+#' surprising changes in neighbouring regions, and joined regions are missing a value if one of the regions they
+#' are made up of is. Replace missing values by zero beforehand if they stand for regions where nothing got counted
+#' @param id name of the column that uniquely identifies the regions, default is "TongfenID"
+#' @param neighbours optional, neighbouring regions as a table with the identifiers of pairs of neighbouring
+#' regions in the first two columns, or as a neighbours list like the ones returned by `spdep::poly2nb`.
+#' By default all regions with intersecting geometries are neighbours, which can miss neighbours
+#' if the geometries have been simplified and don't share their boundaries any more.
+#' @param rel_scale relative decrease that is half way to full surprise, default is `0.25` for a 25\% drop
+#' @param abs_scale absolute decrease that is half way to full surprise, default is `200`
+#' @param p exponent of the norm used to combine the surprises across the timeline into the total surprise.
+#' Large values focus on the most surprising change, 1 adds up the surprises of all changes, default is `4`
+#' @param surprise_cutoff changes with larger surprise count as surprising, default is `0.15`. Only
+#' regions with at least one surprising change are candidates
+#' @param total_surprise_cutoff only regions with larger total surprise are candidates, default is `0.75`
+#' @param cutoff_fact join regions if the total surprise after joining is lower than this share of the
+#' total surprise of the candidate region, default is `0.6`
+#' @param surprise_reduction_const join regions if joining lowers the total surprise by more than this,
+#' default is `0.15`
+#' @param sum_fact join regions if the total surprise after joining is lower than `cutoff_fact * sum_fact`
+#' times the sum of the total surprises of both regions, default is `0.7`
+#' @return A tibble with one row for each candidate region, most surprising first, with the identifier of the
+#' region, the number of surprising changes `surprise_count`, the total surprise `surprise_total`,
+#' the `period` with the most surprising change, the identifier of the `neighbour` that takes away most of
+#' the surprise, the total surprise `surprise_total_joined` after joining both and `join` indicating if both
+#' regions qualify to get joined.
+#' @export
+#'
+#' @examples
+#' # Check 2001 through 2021 dissemination area level population timelines in the
+#' # City of Vancouver for possible geocoding problems
+#' \dontrun{
+#' datasets <- c("CA01","CA06","CA11","CA16","CA21")
+#' meta <- meta_for_additive_variables(datasets,"Population")
+#' data <- get_tongfen_ca_census(regions=list(CSD="5915022"),meta=meta,level="DA",base_geo="CA21")
+#'
+#' anomalies <- tongfen_detect_anomalies(data,paste0("Population_",datasets))
+#' }
+tongfen_detect_anomalies <- function(data,variables,id="TongfenID",neighbours=NULL,
+ rel_scale=0.25,abs_scale=200,p=4,
+ surprise_cutoff=0.15,total_surprise_cutoff=0.75,
+ cutoff_fact=0.6,surprise_reduction_const=0.15,sum_fact=0.7) {
+ params <- anomaly_params(rel_scale,abs_scale,p,surprise_cutoff,total_surprise_cutoff,
+ cutoff_fact,surprise_reduction_const,sum_fact)
+ input <- anomaly_input(data,variables,id,neighbours)
+ candidates <- anomaly_round(input$V,input$edges,params)
+ periods <- paste0(variables[-length(variables)],"-",variables[-1])
+
+ tibble(!!id:=input$id_values[candidates$index],
+ surprise_count=candidates$surprise_count,
+ surprise_total=candidates$surprise_total,
+ period=periods[candidates$period],
+ neighbour=input$id_values[candidates$neighbour],
+ surprise_total_joined=candidates$surprise_total_joined,
+ join=candidates$join)
+}
+
+
+#' Determine regions to join to correct for likely geocoding anomalies
+#'
+#' @description
+#' \lifecycle{experimental}
+#'
+#' Looks for regions with surprising drops in the timeline of a count variable that are complemented by
+#' a neighbouring region, as explained in `tongfen_detect_anomalies`, and joins them. This gets
+#' repeated on the joined regions until there are no more regions left that qualify to get joined.
+#' In each round a region only gets joined with one other region, the most surprising regions go first.
+#'
+#' Joining regions trades geographic detail for timelines that are consistent over time. The parameters
+#' control how aggressively regions get joined and are best calibrated on the data at hand, erring on the side of
+#' joining too few regions risks keeping geocoding problems, erring on the other side risks removing real
+#' changes and needlessly coarsens the geography.
+#'
+#' The result can be used to join the regions via `tongfen_join_regions`, or to update a correspondence
+#' via `tongfen_join_correspondence`.
+#'
+#' @inheritParams tongfen_detect_anomalies
+#' @return A tibble with one row for each region that gets joined with other regions, with the identifier
+#' of the region, the identifier of the joined region it becomes part of in the column named like the
+#' identifier with suffix `_joined`, by default `TongfenID_joined`, and the `round` in which the region
+#' first got joined to another region. The identifier of a joined region is the smallest
+#' identifier of the regions it is made up of.
+#' @export
+#'
+#' @examples
+#' # Correct 2001 through 2021 dissemination area level population timelines in the
+#' # City of Vancouver for likely geocoding problems
+#' \dontrun{
+#' datasets <- c("CA01","CA06","CA11","CA16","CA21")
+#' meta <- meta_for_additive_variables(datasets,"Population")
+#' data <- get_tongfen_ca_census(regions=list(CSD="5915022"),meta=meta,level="DA",base_geo="CA21")
+#'
+#' joins <- tongfen_anomaly_joins(data,paste0("Population_",datasets))
+#' corrected_data <- tongfen_join_regions(data,joins,meta)
+#' }
+tongfen_anomaly_joins <- function(data,variables,id="TongfenID",neighbours=NULL,
+ rel_scale=0.25,abs_scale=200,p=4,
+ surprise_cutoff=0.15,total_surprise_cutoff=0.75,
+ cutoff_fact=0.6,surprise_reduction_const=0.15,sum_fact=0.7) {
+ params <- anomaly_params(rel_scale,abs_scale,p,surprise_cutoff,total_surprise_cutoff,
+ cutoff_fact,surprise_reduction_const,sum_fact)
+ input <- anomaly_input(data,variables,id,neighbours)
+ result <- anomaly_joins(input$V,input$edges,params)
+
+ membership <- result$membership
+ joined <- tabulate(membership)[membership]>1
+ # region with the smallest identifier in each group
+ o <- order(membership,input$ids,method="radix")
+ first <- !duplicated(membership[o])
+ smallest <- integer(max(membership,0L))
+ smallest[membership[o][first]] <- o[first]
+
+ joins <- tibble(!!id:=input$id_values[joined],
+ !!paste0(id,"_joined"):=input$id_values[smallest[membership[joined]]],
+ round=result$round[joined])
+ joins[order(input$ids[smallest[membership[joined]]],input$ids[joined],method="radix"),]
+}
+
+
+# Combine the TongfenUIDs of regions that get joined. A TongfenUID lists the identifiers
+# making up a region as ":,:".
+merge_tongfen_uids <- function(uids) {
+ uids <- unique(uids[!is.na(uids)])
+ # the result must not depend on the order the regions come in
+ uids <- uids[order(uids,method="radix")]
+ parts <- unlist(strsplit(uids," ",fixed=TRUE))
+ parts <- parts[nzchar(parts)]
+ split <- regexpr(":",parts,fixed=TRUE)
+ # not in the expected format, just string them together
+ if (length(parts)==0 || any(split<2)) return(paste0(uids,collapse=" "))
+ columns <- substr(parts,1,split-1)
+ values <- strsplit(substring(parts,split+1),",",fixed=TRUE)
+ vapply(unique(columns),function(column) {
+ v <- unique(unlist(values[columns==column]))
+ paste0(column,":",paste0(v[order(v,method="radix")],collapse=","))
+ },character(1)) %>%
+ paste0(collapse=" ")
+}
+
+# metadata for aggregating data that already has been aggregated to a common geography,
+# where variables are named by their label
+meta_for_joining_regions <- function(data,meta) {
+ meta <- cut_meta(data,meta)
+ # the same variable can be part of several datasets, look up parents within the dataset
+ dataset <- if ("geo_dataset" %in% names(meta)) as.character(meta$geo_dataset) else rep("",nrow(meta))
+ key <- paste(dataset,meta$variable,sep="\x1f")
+ parent <- meta$data_var[match(paste(dataset,meta$parent,sep="\x1f"),key)]
+
+ needs_parent <- meta$rule %in% c("Average","Median","AverageTo") & is.na(parent)
+ if (any(needs_parent))
+ stop(paste0("Can't join regions for ",paste0(unique(meta$data_var[needs_parent]),collapse=", "),
+ ", the data does not have the parent variables needed to aggregate them. ",
+ "Use `tongfen_join_correspondence` and `tongfen_aggregate` on the original data instead."),
+ call.=FALSE)
+
+ meta %>%
+ mutate(variable=.data$data_var,parent=!!parent) %>%
+ select(any_of(c("variable","rule","parent","units","type"))) %>%
+ unique()
+}
+
+
+#' Join regions in data on a common geography
+#'
+#' @description
+#' \lifecycle{experimental}
+#'
+#' Joins regions in data that has already been aggregated to a common geography, for example to correct
+#' for likely geocoding anomalies as determined by `tongfen_anomaly_joins`. The data, and the geometries if
+#' the data is of class sf, of the regions that get joined are aggregated, all other regions are left as they are.
+#'
+#' Variables are aggregated according to the metadata, numeric variables that are not part of the metadata
+#' are assumed to be additive. Variables that are not additive, like averages, can only be aggregated if their
+#' parent variable is part of the data. If that is not the case use `tongfen_join_correspondence` to update the
+#' correspondence the data was built from and aggregate the original data again with `tongfen_aggregate`.
+#'
+#' @param data data on a common geography, with one row per region, for example as returned by
+#' `tongfen_aggregate` or `get_tongfen_ca_census`
+#' @param joins table with the regions to join as returned by `tongfen_anomaly_joins`, with the identifier
+#' of the region and the identifier of the joined region it becomes part of in the column named like the
+#' identifier with suffix `_joined`
+#' @param meta optional metadata containing aggregation rules as for example returned by `meta_for_ca_census_vectors`,
+#' variables are matched by their label. Numeric variables that are not part of the metadata are treated as
+#' additive, if `NULL` (the default) that is the case for all numeric variables
+#' @param id name of the column that uniquely identifies the regions, default is "TongfenID"
+#' @param na.rm logical, determines how NA values should be treated when aggregating variables,
+#' default is `TRUE`
+#' @return The data with the regions joined. Joined regions take the place and the identifier of the
+#' region with the smallest identifier among the regions they are made up of. Variables that are
+#' not numeric and not part of the metadata are `NA` for joined regions.
+#' @export
+#'
+#' @examples
+#' # Correct 2001 through 2021 dissemination area level population timelines in the
+#' # City of Vancouver for likely geocoding problems
+#' \dontrun{
+#' datasets <- c("CA01","CA06","CA11","CA16","CA21")
+#' meta <- meta_for_additive_variables(datasets,"Population")
+#' data <- get_tongfen_ca_census(regions=list(CSD="5915022"),meta=meta,level="DA",base_geo="CA21")
+#'
+#' joins <- tongfen_anomaly_joins(data,paste0("Population_",datasets))
+#' corrected_data <- tongfen_join_regions(data,joins,meta)
+#' }
+tongfen_join_regions <- function(data,joins,meta=NULL,id="TongfenID",na.rm=TRUE) {
+ joined_id <- paste0(id,"_joined")
+ assert(id %in% names(data),paste0("Did not find ",id," in data."))
+ assert(all(c(id,joined_id) %in% names(joins)),
+ paste0("Need joins to have columns ",id," and ",joined_id,"."))
+ data <- data %>% ungroup()
+ ids <- as.character(data[[id]])
+ assert(!anyNA(ids) && !anyDuplicated(ids),
+ paste0("Need ",id," to uniquely identify the regions in data."))
+
+ match_index <- match(ids,as.character(joins[[id]]))
+ listed <- !is.na(match_index)
+ if (!any(listed)) return(data)
+
+ new_id_values <- data[[id]]
+ new_id_values[listed] <- joins[[joined_id]][match_index[listed]]
+ # regions that other regions get joined to are part of the join, even if not listed themselves
+ affected <- listed | ids %in% as.character(new_id_values[listed])
+
+ is_sf <- "sf" %in% class(data)
+ geo_column <- if (is_sf) attr(data,"sf_column") else NULL
+ joined <- data[affected,]
+ joined[[id]] <- new_id_values[affected]
+ new_ids <- as.character(joined[[id]])
+
+ value_columns <- setdiff(names(data),c(id,"TongfenUID",geo_column))
+ join_meta <- tibble(variable=character(0),rule=character(0),parent=character(0),type=character(0))
+ if (!is.null(meta)) join_meta <- meta_for_joining_regions(joined,meta)
+ # results of tongfen calls can have count variables that are not part of the metadata,
+ # like the population, dwelling and household counts for Canadian census data
+ additive_columns <- setdiff(value_columns,join_meta$variable)
+ additive_columns <- additive_columns[vapply(additive_columns,\(v) is.numeric(joined[[v]]),logical(1))]
+ if (length(additive_columns)>0) {
+ message(paste0(ifelse(is.null(meta),"No metadata given, treating all numeric variables as additive: ",
+ "Treating numeric variables that are not part of the metadata as additive: "),
+ paste0(additive_columns,collapse=", ")))
+ join_meta <- bind_rows(join_meta,
+ tibble(variable=additive_columns,rule="Additive",
+ parent=NA_character_,type="Manual"))
+ }
+ dropped_columns <- setdiff(value_columns,join_meta$variable)
+ if (length(dropped_columns)>0)
+ message(paste0("Don't know how to aggregate ",paste0(dropped_columns,collapse=", "),
+ ", setting to NA for joined regions."))
+
+ aggregated <- joined %>%
+ select(all_of(c(id,intersect(join_meta$variable,names(joined)),geo_column))) %>%
+ group_by(!!as.name(id)) %>%
+ aggregate_data_with_meta(join_meta,na.rm=na.rm) %>%
+ ungroup()
+
+ if ("TongfenUID" %in% names(data)) {
+ uids <- vapply(split(as.character(joined$TongfenUID),new_ids),merge_tongfen_uids,character(1))
+ aggregated$TongfenUID <- unname(uids[as.character(aggregated[[id]])])
+ }
+
+ result <- bind_rows(data[!affected,],aggregated)
+ # joined regions take the place of the region they got their identifier from
+ result <- result[order(match(as.character(result[[id]]),ids)),names(data)]
+ if (is_sf) {
+ geometry_types <- unique(as.character(sf::st_geometry_type(result)))
+ if (length(geometry_types)>1 && all(geometry_types %in% c("POLYGON","MULTIPOLYGON")))
+ result <- sf::st_cast(result,"MULTIPOLYGON")
+ }
+ result
+}
+
+
+#' Join regions in a correspondence
+#'
+#' @description
+#' \lifecycle{experimental}
+#'
+#' Updates a correspondence so that the given regions are joined, for example to correct for likely geocoding
+#' anomalies as determined by `tongfen_anomaly_joins`. The updated correspondence can be used in `tongfen_aggregate`
+#' to aggregate data on the coarser common geography, which works for all variables `tongfen_aggregate` can
+#' deal with and for data that was not part of detecting the anomalies.
+#'
+#' @param correspondence correspondence table with columns the unique geographic identifiers for each of the
+#' geographies and the TongfenID and TongfenUID, as for example returned by `estimate_tongfen_correspondence`
+#' or `get_tongfen_correspondence_ca_census`
+#' @param joins table with the regions to join as returned by `tongfen_anomaly_joins`, with columns
+#' `TongfenID` and `TongfenID_joined`
+#' @return The correspondence with updated TongfenID and TongfenUID for the regions that got joined. If the
+#' correspondence has a TongfenMethod column "anomaly" gets added to the method of the regions that got joined.
+#' @export
+#'
+#' @examples
+#' # Correct for likely geocoding problems in dissemination area level population timelines
+#' # and use the updated correspondence to aggregate data on the corrected common geography
+#' \dontrun{
+#' regions <- list(CSD="5915022")
+#' datasets <- c("CA01","CA06","CA11","CA16","CA21")
+#' meta <- meta_for_additive_variables(datasets,"Population")
+#' data <- get_tongfen_ca_census(regions=regions,meta=meta,level="DA",base_geo="CA21")
+#' joins <- tongfen_anomaly_joins(data,paste0("Population_",datasets))
+#'
+#' correspondence <- get_tongfen_correspondence_ca_census(geo_datasets=datasets,
+#' regions=regions,level="DA") %>%
+#' tongfen_join_correspondence(joins)
+#' }
+tongfen_join_correspondence <- function(correspondence,joins) {
+ assert("TongfenID" %in% names(correspondence),"Did not find TongfenID in correspondence.")
+ assert(all(c("TongfenID","TongfenID_joined") %in% names(joins)),
+ "Need joins to have columns TongfenID and TongfenID_joined.")
+ ids <- as.character(correspondence$TongfenID)
+ missing_regions <- setdiff(as.character(joins$TongfenID),ids)
+ if (length(missing_regions)>0)
+ warning(paste0("Did not find ",length(missing_regions)," of the regions to join in the correspondence."))
+
+ match_index <- match(ids,as.character(joins$TongfenID))
+ listed <- !is.na(match_index)
+ if (!any(listed)) return(correspondence)
+
+ ids[listed] <- as.character(joins$TongfenID_joined)[match_index[listed]]
+ # regions that other regions get joined to are part of the join, even if not listed themselves
+ affected <- listed | ids %in% ids[listed]
+ new_ids <- ids[affected]
+ correspondence$TongfenID[affected] <- new_ids
+ if ("TongfenUID" %in% names(correspondence)) {
+ uids <- vapply(split(as.character(correspondence$TongfenUID[affected]),new_ids),
+ merge_tongfen_uids,character(1))
+ correspondence$TongfenUID[affected] <- unname(uids[new_ids])
+ }
+ if ("TongfenMethod" %in% names(correspondence)) {
+ method <- correspondence$TongfenMethod[affected]
+ tagged <- grepl("anomaly",method,fixed=TRUE)
+ method[!tagged] <- paste0(method[!tagged],", anomaly")
+ correspondence$TongfenMethod[affected] <- method
+ }
+ correspondence
+}
+
+#' @import dplyr
+#' @importFrom rlang .data
+NULL
+if(getRversion() >= "2.15.1") utils::globalVariables(c("."))
diff --git a/R/tongfen_ca.R b/R/tongfen_ca.R
index 60caeb6..3d3d491 100644
--- a/R/tongfen_ca.R
+++ b/R/tongfen_ca.R
@@ -1,13 +1,9 @@
-correspondence_ca_census_urls <- list(
- "2006"=list("DB"="https://www12.statcan.gc.ca/census-recensement/2011/geo/ref/files-fichiers/2006_92-156_DB_ID_txt.zip",
- "DA"="https://www12.statcan.gc.ca/census-recensement/2011/geo/ref/files-fichiers/2006_92-156_DA_AD_txt.zip"),
- "2011"=list("DB"="https://www12.statcan.gc.ca/census-recensement/2011/geo/ref/files-fichiers/2011_92-156_DB_ID_txt.zip",
- "DA"="https://www12.statcan.gc.ca/census-recensement/2011/geo/ref/files-fichiers/2011_92-156_DA_AD_txt.zip"),
- "2016"=list("DB"="https://www12.statcan.gc.ca/census-recensement/2011/geo/ref/files-fichiers/2016/2016_92-156_DB_ID_csv.zip",
- "DA"="https://www12.statcan.gc.ca/census-recensement/2011/geo/ref/files-fichiers/2016/2016_92-156_DA_AD_csv.zip"),
- "2021"=list("DB"="https://www12.statcan.gc.ca/census-recensement/2021/geo/aip-pia/correspondence-correspondance/files-fichiers/2021_92-156-X_DB_ID.zip",
- "DA"="https://www12.statcan.gc.ca/census-recensement/2021/geo/aip-pia/correspondence-correspondance/files-fichiers/2021_92-156-X_DA_AD.zip")
-)
+# StatCan correspondence files as parquet, built by data-raw/statcan_correspondence.R
+correspondence_ca_census_url <- function(year,level){
+ base_url <- nullify_blank(getOption("tongfen.statcan_correspondence_url")) %||%
+ "https://mountainmath.s3.ca-central-1.amazonaws.com/tongfen/statcan_correspondence/v1"
+ paste0(sub("/+$","",base_url),"/statcan_correspondence_",year,"_",level,".parquet")
+}
ca_census_base <- c("Population","Dwellings","Households")
@@ -30,15 +26,6 @@ datasets_from_vectors <- function(vs){
ds
}
-GEO_DATASET_LOOKUP <- c(
- setNames(rep("CA1996",1),paste0("TX",seq(2000,2000))),
- setNames(rep("CA01",5),paste0("TX",seq(2001,2005))),
- setNames(rep("CA06",6),paste0("TX",seq(2006,2011))),
- setNames(rep("CA11",4),paste0("TX",seq(2012,2015))),
- setNames(rep("CA16",5),paste0("TX",seq(2016,2020))),
- setNames(rep("CA16",21),paste0("CA",seq(2000,2020),"RMS"))
-)
-
geo_dataset_for_years <- function(years){
require_suggested("cancensus")
dataset_list <- cancensus::list_census_datasets()
@@ -54,7 +41,6 @@ geo_dataset_for_years <- function(years){
geo_dataset_from_dataset <- function(datasets){
require_suggested("cancensus")
- if (TRUE) { # legacy until cancensus updates
datasets <- datasets %>% gsub("^CA11[NF]$","CA11",.) %>% gsub("\\d{4}x","",.)
dataset_list <- cancensus::list_census_datasets()
lapply(datasets, function(ds){
@@ -64,19 +50,9 @@ geo_dataset_from_dataset <- function(datasets){
unique()
}) %>%
unlist()
- } else {
- result <- tibble(dataset=datasets,geo_dataset=GEO_DATASET_LOOKUP[datasets]) %>%
- mutate(geo_dataset=ifelse(is.na(.data$geo_dataset),.data$dataset %>%
- years_from_datasets() %>%
- as.character() %>%
- substr(3,4) %>%
- paste0("CA",.),
- .data$geo_dataset))
- result$geo_dataset
- }
}
-#' Generate metadata from Candian census vectors
+#' Generate metadata from Canadian census vectors
#'
#' @description
#' \lifecycle{maturing}
@@ -92,7 +68,7 @@ geo_dataset_from_dataset <- function(datasets){
#' @examples
#' # Build metadata for vectors
#' \dontrun{
-#' meta <- meta_for_ca_census_vectors("v_CA16_4836","v_CA16_4838","v_CA16_4899")
+#' meta <- meta_for_ca_census_vectors(c("v_CA16_4836","v_CA16_4838","v_CA16_4899"))
#'}
meta_for_ca_census_vectors <- function(vectors){
require_suggested("cancensus")
@@ -104,9 +80,6 @@ meta_for_ca_census_vectors <- function(vectors){
nn[nn==""]=vectors[nn==""]
}
- if (length(vectors)==0) {
- meta <- tibble::tibble(variable=NA,label=NA,dataset=datasets_from_vectors(vectors))
- }
meta <- tibble::tibble(variable=vectors,label=nn,dataset=datasets_from_vectors(vectors)) %>%
mutate(type="Original", aggregation="0",units=NA)
datasets <- meta$dataset %>%
@@ -138,13 +111,13 @@ meta_for_ca_census_vectors <- function(vectors){
select(variable="parent","dataset") %>%
mutate(type="Extra",aggregation="Additive",rule="Additive") %>%
filter(!is.na(.data$variable),!.data$variable %in% meta$variable) %>%
- filter(!duplicated(.data$variable,.data$dataset)) %>%
+ distinct(.data$variable,.data$dataset,.keep_all=TRUE) %>%
mutate(label=.data$variable)
if (nrow(extras)>0) {
meta <- meta %>%
bind_rows(extras) %>%
- filter(!duplicated(.data$variable,.data$dataset))
+ distinct(.data$variable,.data$dataset,.keep_all=TRUE)
}
meta <- meta %>%
@@ -155,13 +128,13 @@ meta_for_ca_census_vectors <- function(vectors){
-#' Generate metadata from Candian census vectors
+#' Generate metadata from Canadian census vectors
#'
#' @description
#' \lifecycle{maturing}
#'
#' Add Population, Dwellings, and Household counts to metadata
-#' @param meta ribble with metadata as for example provided by `meta_for_ca_census_vectors`
+#' @param meta tibble with metadata as for example provided by `meta_for_ca_census_vectors`
#' @return tibble with metadata
add_census_ca_base_variables <- function(meta){
new_meta <- meta$geo_dataset %>%
@@ -184,6 +157,11 @@ add_census_ca_base_variables <- function(meta){
#' @description
#' \lifecycle{maturing}
#'
+#' The correspondence files are downloaded from a mirror of the Statistics Canada correspondence files
+#' and cached in the tongfen cache directory. The cached files are checked against the mirror once per session
+#' and get downloaded again if they changed. The location of the mirror can be changed via the
+#' `tongfen.statcan_correspondence_url` option.
+#'
#' @param year census year, only 2006 through 2021 are supported
#' @param level geographic level, DA or DB
#' @param refresh reload the correspondence files, default is `FALSE`
@@ -193,36 +171,9 @@ get_single_correspondence_ca_census_for <- function(year,level=c("DA","DB"),refr
year=as.character(year)[1]
if (!(level %in% c("DA","DB"))) stop("Level needs to be DA or DB")
if (!(year %in% c("2006","2011","2016","2021"))) stop("Year needs to be 2006, 2011, 2016, or 2021")
- new_field=paste0(level,"UID",year)
- old_field=paste0(level,"UID",as.integer(year)-5)
- path=file.path(tongfen_cache_dir(),paste0("statcan_correspondence_",year,"_",level,".csv"))
- if (refresh || !file.exists(path)) {
- url=correspondence_ca_census_urls[[year]][[level]]
- tmp=tempfile()
- utils::download.file(url,tmp)
- exdir=file.path(tempdir(),paste0("correspondence_",year,"_",level))
- if (dir.exists(exdir)) unlink(exdir,recursive=TRUE)
- dir.create(exdir,showWarnings = FALSE)
- utils::unzip(tmp,exdir=exdir)
- file=dir(exdir,"\\.txt|\\.csv")
- if (length(file)==0) {
- p<-dir(exdir)[1]
- if (dir.exists(file.path(exdir,p))) {
- exdir=file.path(exdir,p)
- file=dir(exdir,"\\.txt|\\.csv")
- }
- }
- if (level=="DB") headers=c(new_field,old_field,"flag") else headers=c(new_field,old_field,paste0("DBUID",year),"flag")
- unwanted <- paste0(level,"UID",year)
- d<-readr::read_csv(file.path(exdir,file),col_types = readr::cols(.default = "c"),col_names = headers) %>%
- select(all_of(c(new_field,old_field,"flag"))) %>%
- unique() %>%
- filter(!grepl(unwanted,!!as.name(new_field)))
- readr::write_csv(d,path)
- unlink(tmp)
- unlink(exdir,recursive = TRUE)
- }
- result <- readr::read_csv(path,col_types = readr::cols(.default = "c"))
+ path=file.path(tongfen_cache_dir(),paste0("statcan_correspondence_",year,"_",level,".parquet"))
+ cached_download(correspondence_ca_census_url(year,level),path,refresh=refresh)
+ result <- tibble::as_tibble(nanoparquet::read_parquet(path))
# manual corrections
if (year=="2021" && level=="DB") {
@@ -243,7 +194,7 @@ get_single_correspondence_ca_census_for <- function(year,level=c("DA","DB"),refr
#' @description
#' \lifecycle{maturing}
#'
-#' Get correspondence file for several Candian censuses on a common geography. Requires sf and cancensus package to be available
+#' Get correspondence file for several Canadian censuses on a common geography. Requires sf and cancensus package to be available
#'
#' @param regions census region list, should be inclusive list of GeoUIDs across censuses
#' @param geo_datasets vector of census geography dataset identifiers
@@ -253,7 +204,7 @@ get_single_correspondence_ca_census_for <- function(year,level=c("DA","DB"),refr
#' this method only works for "DB", "DA" and "CT" levels.
#' * "estimate" uses `estimate_tongfen_correspondence` to build up the common geography from scratch based on geographies.
#' * "identifier" assumes regions with identical geographic identifier are identical, and builds up the the correspondence for regions with unmatched geographic identifiers.
-#' @param tolerance tolerance for `estimate_tongen_correspondence` in metres, default value is 50 metres,
+#' @param tolerance tolerance for `estimate_tongfen_correspondence` in metres, default value is 50 metres,
#' only used when method is 'estimate' or 'identifier'
#' @param quiet suppress download progress output, default is `FALSE`
#' @param refresh optional character, refresh data cache for this call, (default `FALSE`)
@@ -352,7 +303,7 @@ get_tongfen_correspondence_ca_census <- function(geo_datasets, regions, level="C
correspondence_years=all_geo_years[-1]
correspondence <- correspondence_years %>%
lapply(function(year){
- c <- get_single_correspondence_ca_census_for(year,statcan_level) %>%
+ c <- get_single_correspondence_ca_census_for(year,statcan_level,refresh=refresh) %>%
select(-"flag")
previous_year <- all_geo_years[which(all_geo_years==year)-1]
ds1 <- all_geo_datasets[all_geo_years==year]
@@ -396,15 +347,15 @@ get_tongfen_correspondence_ca_census <- function(geo_datasets, regions, level="C
}
-#' Togfen data from several Canadian censuses
+#' Tongfen data from several Canadian censuses
#'
#' @description
#' \lifecycle{maturing}
#'
-#' Get data from several Candian censuses on a common geography. Requires sf and cancensus package to be available
+#' Get data from several Canadian censuses on a common geography. Requires sf and cancensus package to be available
#'
#' @param regions census region list, should be inclusive list of GeoUIDs across censuses
-#' @param meta metadata for the census veraiables to aggregate, for example as returned
+#' @param meta metadata for the census variables to aggregate, for example as returned
#' by \code{meta_for_ca_census_vectors}.
#' @param level aggregation level to return data on (default is "CT")
#' @param method tongfen method, options are "statcan" (the default), "estimate", "identifier".
@@ -416,7 +367,7 @@ get_tongfen_correspondence_ca_census <- function(geo_datasets, regions, level="C
#' any geographic data
#' @param na.rm logical, determines how NA values should be treated when aggregating variables,
#' default is `FALSE`
-#' @param tolerance tolerance for `estimate_tongen_correspondence` in metres, default value is 50 metres,
+#' @param tolerance tolerance for `estimate_tongfen_correspondence` in metres, default value is 50 metres,
#' only used when method is 'estimate' or 'identifier'
#' @param quiet suppress download progress output, default is `FALSE`
#' @param refresh optional character, refresh data cache for this call, (default `FALSE`)
diff --git a/R/tongfen_ca_deprecated.R b/R/tongfen_ca_deprecated.R
index 418ff0e..f0cf13b 100644
--- a/R/tongfen_ca_deprecated.R
+++ b/R/tongfen_ca_deprecated.R
@@ -85,7 +85,7 @@ get_tongfen_census_da <- function(regions,vectors,geo_format=NA,use_cache=TRUE,n
#' @return dataframe with variables on common geography
#' @export
get_tongfen_ca_census_ct_from_da <- function(regions,vectors,geo_format=NA,use_cache=TRUE,na.rm=TRUE,quiet=TRUE) {
- lifecycle::deprecate_warn("0.2.0", "get_tongfen_census_da()", "get_tongfen_census_ca()")
+ lifecycle::deprecate_warn("0.2.0", "get_tongfen_ca_census_ct_from_da()", "get_tongfen_ca_census()")
meta <- meta_for_ca_census_vectors(vectors)
base_geo <- if (is.na(geo_format)) NULL else meta$geo_dataset %>% unique() %>% sort() %>% first()
@@ -106,11 +106,13 @@ get_tongfen_ca_census_ct_from_da <- function(regions,vectors,geo_format=NA,use_c
#' \lifecycle{deprecated}
#'
#' Aggregate variables to common CTs, returns data2 on new tiling matching data1 geography
-#' @param data1 cancensus CT level datatset for year1 < year2 to serve as base for common geography
-#' @param data2 cancensus CT level datatset for year2 to be aggregated to common geography
+#' @param data1 cancensus CT level dataset for year1 < year2 to serve as base for common geography
+#' @param data2 cancensus CT level dataset for year2 to be aggregated to common geography
#' @param data2_sum_vars vector of variable names to by summed up when aggregating geographies
#' @param data2_group_vars optional vector of grouping variables
#' @param na.rm optional parameter to remove NA values when summing, default = `TRUE`
+#' @return `data2` with the variables in `data2_sum_vars` aggregated to a common geography matching `data1`,
+#' identified by the `GeoUID` of `data1`
#' @export
tongfen_ca_census_ct <- function(data1,data2,data2_sum_vars,data2_group_vars=c(),na.rm=TRUE) {
lifecycle::deprecate_warn("0.2.0", "tongfen_ca_census_ct()", "tongfen_aggregate()")
@@ -159,7 +161,7 @@ tongfen_ca_census_ct <- function(data1,data2,data2_sum_vars,data2_group_vars=c()
#'
#' @description
#' \lifecycle{deprecated}
-#' Joins the StatCan correspodence files for several census years
+#' Joins the StatCan correspondence files for several census years
#'
#' @param years list of census years
#' @param level geographic level, DA or DB
diff --git a/R/tongfen_ca_estimate.R b/R/tongfen_ca_estimate.R
index 6fcaecb..bf24258 100644
--- a/R/tongfen_ca_estimate.R
+++ b/R/tongfen_ca_estimate.R
@@ -25,11 +25,12 @@
#' one of these for all variables.
#' @param na.rm how to deal with NA values, default is \code{FALSE}.
#' @param quiet suppress progress messages
+#' @return `geometry` with the estimated values for the census variables specified by `meta`
#' @export
#'
#' @examples
-#' # Estimate a common geography for 2006 and 2016 dissemination areas in the City of Vancouver
-#' # based on the geographic data and check estimation errors
+#' # Estimate the 2016 population within 1 km of Toronto City Hall from dissemination area level
+#' # census data
#' \dontrun{
#' toronto_city_hall <- sf::st_point(c(-79.3839,43.6534)) %>%
#' sf::st_sfc(crs=4326) %>%
@@ -42,7 +43,7 @@
#' data <- tongfen_estimate_ca_census(toronto_city_hall,meta,level="DA",intersection_level="CT")
#'
#' print(paste0("Approximately ",scales::comma(data$Population,accuracy=100),
-#' " people live within a 1 km radius of Toronto City."))
+#' " people live within a 1 km radius of Toronto City Hall."))
#'
#'}
tongfen_estimate_ca_census <- function(geometry, meta, level,
@@ -85,7 +86,7 @@ tongfen_estimate_ca_census <- function(geometry, meta, level,
census_data <- g
}
- result <- tongfen_estimate(target = geometry, source = census_data, meta = meta,na.rm = na.rm)
+ tongfen_estimate(target = geometry, source = census_data, meta = meta,na.rm = na.rm)
}
diff --git a/R/tongfen_estimate.R b/R/tongfen_estimate.R
index 8342e84..75f63ce 100644
--- a/R/tongfen_estimate.R
+++ b/R/tongfen_estimate.R
@@ -13,7 +13,9 @@
#' @param meta metadata for variable aggregation, see `meta_for_additive_variables` and `meta_for_ca_census_vectors` for more information
#' on how to construct metadata.
#' @param na.rm remove NA values when aggregating, default is FALSE
-#' @return `target` with estimated quantities from `source` as specified by `meta`
+#' @return `target` with estimated quantities from `source` as specified by `meta`, regions in `target`
+#' that don't overlap with `source` have `NA` values. Columns in `target` can't have the same name
+#' as the variables to be estimated.
#' @export
#'
#' @examples
@@ -46,6 +48,12 @@ tongfen_estimate <- function(target,source,meta,na.rm=FALSE) {
cut_meta(source,.) %>%
mutate(var_name=paste0("v",row_number()))
+ clashes <- intersect(meta$data_var,names(target))
+ if (length(clashes)>0) {
+ stop(paste0("Target already has columns named ",paste0(clashes,collapse=", "),
+ ", please rename or remove these before estimating."))
+ }
+
# rename variables, st_interpolate_aw does not handle column names with special characters
safe_rename_vars <- setNames(meta$data_var,meta$var_name)
safe_rename_back <- setNames(meta$var_name,meta$data_var)
@@ -54,30 +62,33 @@ tongfen_estimate <- function(target,source,meta,na.rm=FALSE) {
i = suppressMessages(st_intersection(st_geometry(source), st_geometry(target)))
idx = attr(i, "idx")
- gc = which(st_is(i, "GEOMETRYCOLLECTION"))
- i[gc] = st_collection_extract(i[gc], "POLYGON")
+ # the area only counts the polygonal parts of intersections, source regions that only
+ # touch a target region along a boundary don't overlap and must not contribute
+ area_st <- as.numeric(st_area(i))
+ overlaps <- which(area_st > 0)
+ idx <- idx[overlaps,,drop=FALSE]
+ area_st <- area_st[overlaps]
source <- source %>% rename(!!!safe_rename_vars)
- source_area <- unclass(st_area(source))
+ source_area <- as.numeric(st_area(source))
x_st <- source[idx[,1],, drop=FALSE] %>%
+ st_drop_geometry() %>%
select(names(safe_rename_vars)) %>%
- pre_scale(meta,meta_var = "var_name") %>%
- mutate(...area_st = st_area(i) %>% unclass,
- ...area_s = source_area[idx[,1]]) %>%
- mutate(...factor = .data$...area_st/.data$...area_s) %>%
- mutate(...partial = .data$...factor < 0.99) %>%
- st_drop_geometry()
+ pre_scale(meta,meta_var = "var_name")
- x_st[meta$var_name] <- lapply(x_st[meta$var_name], `*`, x_st$...factor)
+ # variables to sum up, including the weights for averages added when pre-scaling
+ sum_vars <- names(x_st)
+ x_st[sum_vars] <- lapply(x_st[sum_vars], `*`, area_st/source_area[idx[,1]])
- x_st <- stats::aggregate(x_st, list(idx[,2]), sum, na.rm=na.rm)
+ # target regions without overlap with the source don't show up here and end up as NA
+ x_st <- x_st %>%
+ mutate(!!unique_key:=idx[,2]) %>%
+ group_by(.data[[unique_key]]) %>%
+ summarize(across(all_of(sum_vars), \(x) sum(x,na.rm=na.rm)),.groups="drop")
result <- target %>%
- left_join(x_st %>%
- select(-all_of(c("...factor", "...partial", "...area_s", "...area_st"))) %>%
- rename(!!unique_key:="Group.1"),
- by=unique_key) %>%
+ left_join(x_st,by=unique_key) %>%
select(-all_of(unique_key)) %>%
post_scale(meta,meta_var = "var_name") %>%
rename(!!!safe_rename_back)
@@ -117,17 +128,17 @@ tongfen_estimate <- function(target,source,meta,na.rm=FALSE) {
#' @param target custom geography
#' @param source input geography
#' @param target_id name of the column in `target` table with unique id (character)
-#' @return `source` with extra column with name `"target_id"` and column `...overlap_fraction` with
-#' the proportion of overlap of the target geometry with the respective `target_id`
+#' @return `source` with extra column with the name given by `target_id` and column `...overlap_fraction` with
+#' the proportion of the area of the source region that overlaps with the region in `target` with that id
#' @export
#'
#' @examples
-#' # Estimate 2006 Populatino in the City of Vancouver dissemination ares on 2016 census geoographies
+#' # Tag 2016 dissemination areas in the City of Vancouver by the 2006 census tract they overlap
+#' # the most with
#' \dontrun{
-#' geo1 <- cancensus::get_census("CA06",regions=list(CSD="5915022"),geo_format='sf',level='DA')
+#' geo1 <- cancensus::get_census("CA06",regions=list(CSD="5915022"),geo_format='sf',level='CT')
#' geo2 <- cancensus::get_census("CA16",regions=list(CSD="5915022"),geo_format='sf',level='DA')
-#' meta <- meta_for_additive_variables("CA06","Population")
-#' result <- tongfen_estimate(geo2 %>% rename(Population_2016=Population),geo1,meta)
+#' result <- tongfen_tag_largest_overlap(geo2,geo1 %>% select(CT_2006=GeoUID),"CT_2006")
#'}
tongfen_tag_largest_overlap <- function(source, target, target_id) {
target_geo_types <- target %>% sf::st_geometry_type() %>% unique
diff --git a/R/tongfen_us.R b/R/tongfen_us.R
index 19ecf35..7e9790a 100644
--- a/R/tongfen_us.R
+++ b/R/tongfen_us.R
@@ -59,14 +59,9 @@ get_us_ct_correspondence_path <- function(state,year){
# the blocks they are made up of
get_us_ct_correspondence_2020 <- function(state,min_area_share=0.01,
cache_path=getOption("tongfen.cache_path")) {
- cache_path = file.path(cache_path %||% tempdir(),"us_data")
-
path <- get_us_ct_correspondence_path(state,2020)
- local_path <- file.path(cache_path,basename(path))
- if (!file.exists(local_path)) {
- if (!dir.exists(cache_path)) dir.create(cache_path, recursive = TRUE)
- utils::download.file(path,local_path,quiet = TRUE)
- }
+ local_path <- file.path(us_cache_dir(cache_path),basename(path))
+ cached_download(path,local_path,check_remote=FALSE)
blocks <- readr::read_delim(local_path,delim="|",progress=FALSE,
col_types=readr::cols_only(
STATE_2010="c",COUNTY_2010="c",TRACT_2010="c",BLK_2010="c",
@@ -97,13 +92,8 @@ get_us_ct_correspondence_2020 <- function(state,min_area_share=0.01,
get_us_ct_correspondence_2010 <- function(state,min_area_share=0.01,
cache_path=getOption("tongfen.cache_path")){
path <- get_us_ct_correspondence_path(state,"2010")
- file <- basename(path)
- cache_path = file.path(cache_path %||% tempdir(),"us_data")
- local_path <- file.path(cache_path,file)
- if (!file.exists(local_path)) {
- if (!dir.exists(cache_path)) dir.create(cache_path, recursive = TRUE)
- utils::download.file(path,local_path,quiet=TRUE)
- }
+ local_path <- file.path(us_cache_dir(cache_path),basename(path))
+ cached_download(path,local_path,check_remote=FALSE)
d<-readr::read_csv(local_path,progress=FALSE,
col_names=c("STATE00","COUNTY00","TRACT00","GEOID00",
"POP00","HU00","PART00","AREA00","AREALAND00",
@@ -130,12 +120,8 @@ get_us_ct_correspondence_2010 <- function(state,min_area_share=0.01,
get_us_ct_correspondence_2000 <- function(state,min_area_share=0.01,
cache_path=getOption("tongfen.cache_path")){
path <- get_us_ct_correspondence_path(state,"2000")
- cache_path = file.path(cache_path %||% tempdir(),"us_data")
- local_path <- file.path(cache_path,basename(path))
- if (!file.exists(local_path)) {
- if (!dir.exists(cache_path)) dir.create(cache_path, recursive = TRUE)
- utils::download.file(path,local_path,quiet=TRUE)
- }
+ local_path <- file.path(us_cache_dir(cache_path),basename(path))
+ cached_download(path,local_path,check_remote=FALSE)
d <- readr::read_fwf(local_path,
readr::fwf_cols(STATE90=c(1,2),COUNTY90=c(3,5),TRACT90BASE=c(6,9),
TRACT90SUF=c(10,11),PART90=c(12,12),POP90TRACT=c(13,21),
@@ -202,15 +188,21 @@ get_us_ct_correspondence <- function(state, datasets, min_area_share=0.01,
# the 2000 to 2010 county subdivision comparability file, covering all states
get_us_county_subdivision_correspondence <- function(cache_path=getOption("tongfen.cache_path")){
require_suggested("readxl")
- cache_path = file.path(cache_path %||% tempdir(),"us_data")
+ cache_path = us_cache_dir(cache_path)
file <- "Cousub_comparability.xlsx"
local_path <- file.path(cache_path,file)
if (!file.exists(local_path)) {
if (!dir.exists(cache_path)) dir.create(cache_path, recursive = TRUE)
- tmp=tempfile(fileext = ".zip")
+ # download and unpack in a temporary directory, an interrupted download must not leave
+ # a broken file in the cache
+ tmp_dir <- tempfile()
+ dir.create(tmp_dir)
+ on.exit(unlink(tmp_dir,recursive=TRUE))
+ tmp <- file.path(tmp_dir,"cousub_comparabilityxls.zip")
path="https://www2.census.gov/geo/docs/maps-data/data/comp/cousub_comparabilityxls.zip"
- utils::download.file(path,tmp,quiet=TRUE)
- utils::unzip(tmp,exdir = cache_path)
+ utils::download.file(path,tmp,mode="wb",quiet=TRUE)
+ utils::unzip(tmp,files=file,exdir=tmp_dir)
+ file.copy(file.path(tmp_dir,file),local_path,overwrite=TRUE)
}
readxl::read_xlsx(local_path)
}
@@ -223,14 +215,10 @@ get_us_county_subdivision_correspondence <- function(cache_path=getOption("tongf
# the share is taken over the larger of the two.
get_us_county_subdivision_correspondence_2020 <- function(min_area_share=0.01,
cache_path=getOption("tongfen.cache_path")){
- cache_path = file.path(cache_path %||% tempdir(),"us_data")
path <- paste0("https://www2.census.gov/geo/docs/maps-data/data/rel2020/cousub/",
"tab20_cousub20_cousub10_natl.txt")
- local_path <- file.path(cache_path,basename(path))
- if (!file.exists(local_path)) {
- if (!dir.exists(cache_path)) dir.create(cache_path, recursive = TRUE)
- utils::download.file(path,local_path,quiet=TRUE)
- }
+ local_path <- file.path(us_cache_dir(cache_path),basename(path))
+ cached_download(path,local_path,check_remote=FALSE)
d <- readr::read_delim(local_path,delim="|",progress=FALSE,
col_types=readr::cols_only(GEOID_COUSUB_10="c",GEOID_COUSUB_20="c",
AREALAND_COUSUB_10="d",AREAWATER_COUSUB_10="d",
@@ -300,7 +288,8 @@ get_us_county_subdivision_correspondence_for <- function(state, datasets, min_ar
#' geographies at the risk of separating regions that did change. No region is ever dropped,
#' if all of its parts are slivers its largest part is kept.
#' @param cache_path optional path to cache the relationship files in, defaults to the
-#' `tongfen.cache_path` option and falls back to a temporary directory
+#' `tongfen.cache_path` option. If that is not set the `tongfen.cache_path` environment variable
+#' and the `custom_data_path` option are used, falling back to a temporary directory
#' @return tibble with one row per census geography, a GEOID column for each requested census,
#' and the common geography identified by `TongfenID` and `TongfenUID`.
#' @export
@@ -351,7 +340,7 @@ sumfile_for_dataset <- function(sumfile, ds){
unname(sumfile[[ds]])
}
-#' Get US census data for 2000 and 2010 census on common census tract based geography
+#' Get US census data for several censuses on a common geography
#'
#' @description
#' \lifecycle{maturing}
@@ -369,14 +358,15 @@ sumfile_for_dataset <- function(sumfile, ds){
#' @param meta metadata for variables to retrieve
#' @param level aggregation level to return the data on. At this stage, the only valid levels are 'tract' and 'county subdivision'.
#' @param survey survey to get data for, supported options is "census"
-#' @param base_geo census year to use as base geography, default is `2010`.
+#' @param base_geo dataset to use as base geography, for example `"dec2010"`, has to be one of
+#' the datasets in `meta`. Default is `NULL`, which uses the first dataset in `meta`.
#' @param min_area_share minimum share of area two geographies have to have in common to count
#' as related, default is `0.01`, see \code{\link{get_tongfen_correspondence_us_census}}.
#' @param sumfile summary file to read the variables from, either a single value used for all
#' censuses or a vector named by dataset, for example `c(dec2010="sf1", dec2020="dhc")`. Default
#' is `NULL`, which leaves the choice to tidycensus. Note that tidycensus defaults the 2020
#' census to the PL 94-171 redistricting file, most 2020 variables need `sumfile="dhc"`.
-#' @return sf object with (wide form) census variables with census year as suffix (separated by underdcore "_").
+#' @return sf object with (wide form) census variables with census year as suffix (separated by underscore "_").
#' @export
#'
#' @examples
@@ -404,7 +394,7 @@ get_tongfen_us_census <- function(regions,meta,level='tract',survey="census",
assert(base_geo %in% datasets,paste0("base_geo has to be one of the datasets ",paste0(datasets,collapse=", ")))
invalid_datasets <- setdiff(datasets,names(valid_us_census_datasets))
assert(length(invalid_datasets)==0, paste0("Invalid datasets :",paste0(invalid_datasets,collapse = ", ")))
- assert(level %in% c('tract','county subdivision'),"Only census tracts and counties are supported right now.")
+ assert(level %in% c('tract','county subdivision'),"Only census tracts and county subdivisions are supported right now.")
assert(survey %in% c('census'),"Only census surveys are supported right now.")
if (!is.null(sumfile)) {
if (is.null(names(sumfile))) {
diff --git a/README.md b/README.md
index 7a1ad3d..d86e2eb 100644
--- a/README.md
+++ b/README.md
@@ -29,13 +29,13 @@ library(tongfen)
### Caching correspondence files
The `get_tongfen_ca_census` and `get_tongfen_correspondence_ca_census` methods make use of the StatCan correspondence
-files when run with `method = "statcan"`. To speed up this process it is useful to permanently cache these files instead of having to download them repeatedly. If caching is desired, set either
+files when run with `method = "statcan"`. Statistics Canada no longer allows programmatic downloads of these files, so the package downloads them from a mirror hosting the files in parquet format. To speed up this process it is useful to permanently cache these files instead of having to download them again in every session. If caching is desired, set either
* `options("tongfen.cache_path"="")`
* `Sys.setenv("tongfen.cache_path"="")`
* `options("custom_data_path"="")`
-in your `.Rprofile` or `.Renviron` file.
+in your `.Rprofile` or `.Renviron` file. Cached files are checked against the mirror once per session and only get downloaded again if they changed, when offline the cached files are used as they are.
## General TongFen
@@ -52,7 +52,7 @@ A convenience function to validate geographic TongFen fit via area comparison is
Finding a common tiling of several different yet congruent geographies is only one part of the problem TongFen addresses, aggregating up the variables is the other part. The `tongfen` package deals with this using a *metadata* table that specifies how variables should be aggregated. In it's simplest form values are simply added up. The `meta_for_additive_variables` convenience function builds the metadata for additive variables. Metadata for non-additive variables like averages, ratios or percentages needs more care to build, it requires additional information on the **parent variable** that specifies the denominator of the average, ratio or percentage. Other data, like medians, can't be aggregated up, although `tongfen` can provide estimates of medians on aggregated geographies by treating them as averages.
### Packaged data
-The package ships with a subset of [voting data from Elections Canada](https://www.elections.ca/content.aspx?section=ele&dir=pas&document=index&lang=e) for the 42nd and 43rd federal elections as well as the polling district geographies for the [42nd](https://open.canada.ca/data/en/dataset/6a78ccfd-6bba-4109-b040-87cb8c71ec35) and [43rd](https://open.canada.ca/data/en/dataset/e70e3263-8584-4f22-94cb-8c15b616cbfc). This facilitates running the example vignette on polling districts without having to download external data. Both are available as open data covered under the [Open Government Licence - Canda](https://open.canada.ca/en/open-government-licence-canada).
+The package ships with a subset of [voting data from Elections Canada](https://www.elections.ca/content.aspx?section=ele&dir=pas&document=index&lang=e) for the 42nd and 43rd federal elections as well as the polling district geographies for the [42nd](https://open.canada.ca/data/en/dataset/6a78ccfd-6bba-4109-b040-87cb8c71ec35) and [43rd](https://open.canada.ca/data/en/dataset/e70e3263-8584-4f22-94cb-8c15b616cbfc). This facilitates running the example vignette on polling districts without having to download external data. Both are available as open data covered under the [Open Government Licence - Canada](https://open.canada.ca/en/open-government-licence-canada).
## Data-specific implementations
The need for TongFen comes up frequently with certain types of geographies. Census geographies is one such example. In some cases these data sources come with their own correspondence files that go beyond geographic matchup but also join regions to alleviate data integrity problems like geocoding issues.
@@ -67,14 +67,15 @@ The package is well-integrated to work with Canadian census data in two essentia
### US census data
-* `get_tongfen_us_census` integrates the data acquisition (via the [**tidycensus** package](https://walker-data.com/tidycensus/index.html)) with TongFen, and adds the tongfen `method = "census.gov"` to use the US Census Bureau correspondence files for matching.
+* `get_tongfen_us_census` integrates the data acquisition (via the [**tidycensus** package](https://walker-data.com/tidycensus/index.html)) with TongFen, using the US Census Bureau relationship files to build the common geography.
+* `get_tongfen_correspondence_us_census` breaks out the correspondence generation from the US Census Bureau relationship files, to tongfen data that comes on census geographies but is obtained by other means. The relationship files are cached in the `us_data` folder of the cache path described above.
## Other implementations
The `tongfen` package is open to add extensions for other specialized data sources, as well as extensions of existing ones.
## Fixed target geography estimation
-When geographies aren't sufficiently congruent or the target geography is fixed, we won't be able to use the `tongfen` methods to compute the data on a common geography but have to instead rely on estimates. The `tongfen_estimate` makes no assumption on the underlying geographies and returns estimates of the data on the target geography. It uses area-weighted interpolation to achieve this, and can be refined to dasymmetric estimates using the `proportional_reaggregate` function.
+When geographies aren't sufficiently congruent or the target geography is fixed, we won't be able to use the `tongfen` methods to compute the data on a common geography but have to instead rely on estimates. The `tongfen_estimate` makes no assumption on the underlying geographies and returns estimates of the data on the target geography. It uses area-weighted interpolation to achieve this, and can be refined to dasymetric estimates using the `proportional_reaggregate` function.
This method has the example that it works independent of the nature of the underlying geographies, but comes at the heavy price of only being an estimate. To be useful for research purposes we also need methods to estimate the errors this introduces and the effects this has on subsequent analysis results.
@@ -84,8 +85,8 @@ Methods to facilitate this are still under active development.
If you wish to cite tongfen:
- von Bergmann, J. (2024). tongfen: R package to
- Make Data Based on Different Geographies Comparable. v0.3.7.
+ von Bergmann, J. (2026). tongfen: R package to
+ Make Data Based on Different Geographies Comparable. v0.3.9.
DOI: 10.32614/CRAN.package.tongfen
@@ -94,9 +95,9 @@ A BibTeX entry for LaTeX users is
@Manual{tongfen,
author = {Jens {von Bergmann}},
title = {tongfen: R package to Make Data Based on Different Geographies Comparable},
- year = {2024},
+ year = {2026},
doi = {10.32614/CRAN.package.tongfen},
- note = {R package version 0.3.7},
+ note = {R package version 0.3.9},
url = {https://mountainmath.github.io/tongfen/},
}
```
diff --git a/cran-comments.md b/cran-comments.md
index c6fa36b..298facd 100644
--- a/cran-comments.md
+++ b/cran-comments.md
@@ -1,63 +1,29 @@
-# tongfen v.0.3.8
-## Breaking changes
-- `get_tongfen_ca_census` now honours its `base_geo`, `na.rm`, `tolerance`, `crs` and
- `data_transform` arguments, all of which were silently ignored
-- removed the `area_mismatch_cutoff` argument from `get_tongfen_ca_census` and
- `get_tongfen_correspondence_ca_census`, it never had any effect
+# tongfen v.0.3.9
## Major changes
-- correspondence tables are now built via a vectorised connected components pass instead of
- a row-by-row union-find, making tongfen on large geographies dramatically faster
-- the "statcan" method no longer downloads census geometries it does not use
-- dissolving geometries skips regions that don't need to be merged
-- new `get_tongfen_correspondence_us_census`, US tract correspondence tables now reach back to
- the 1990 census and county subdivisions forward to the 2020 census
-- US correspondence tables no longer chain regions together over slivers, and no longer strip
- leading zeros off 2020 census tract identifiers
+- new experimental functions `tongfen_detect_anomalies`, `tongfen_anomaly_joins`, `tongfen_join_regions`
+ and `tongfen_join_correspondence` to detect and correct for likely geocoding anomalies in timelines
+ on a common geography, together with a new vignette
+- StatCan correspondence files are now downloaded as parquet files from a mirror, Statistics Canada
+ put the original files behind a browser check that blocks programmatic downloads, which broke
+ `method = "statcan"`
## Minor changes
-- `get_tongfen_us_census` gained a `sumfile` argument, passed through to tidycensus
-- `get_tongfen_correspondence_ca_census` gained a `crs` argument for the spatial intersections
-- missing geographic identifiers no longer merge unrelated regions into one common geography
-- fix crash when tongfen-ing census tracts across non-adjacent censuses
-- US county subdivision data errors out up front on censuses it can't be matched across
-- several fixes to the deprecated `get_tongfen_census_*` functions
-- fix `proportional_reaggregate` ignoring all but the first base variable when `base` names a
- different variable per category
-- fix `estimate_tongfen_correspondence` with `method="identifier"` erroring out when every
- geographic identifier matches
-- packages in Suggests (`cancensus`, `tidycensus`, `readxl`) are now used conditionally, with an
- actionable message when they are not installed
-- faster `check_tongfen_areas` and `aggregate_correspondences`
-
-# tongfen v.0.3.7
-## Major changes
-- accommodate factors in proportional_reaggregate
-- sizable performance increases
-- squish several edge case bugs
-
-# tongfen v.0.3.6
-## Major changs
-- better downsampling that can also accommodate averages
-- performance improvements
-## Minor changes
-- better documentation
-- allow for datasets vartiables by census year for canadian data
-- fix issue where some metadata might get duplicated
-
-
-# Update v.0.3.3
-- Fix compatibility issue with changes in {sf} package
-- More reliable GitHub action CRAN checks
-
-# Update v.0.3.2
-- Added `tongfen_estimate_ca_census` function for new CensusMapper endpoint, tying into new {cancensus} functionality.
-- Custom impelementation of `tongfen_etimate` for finer control
-- Fix compatibility issue with changes in {sf} package
-
-# Submission - v.0.3
+- combining correspondences across three or more datasets no longer runs into a cross join
+- fix `proportional_reaggregate` giving wrong results when the finer level data already has values
+- fix `tongfen_estimate` mishandling intersections that are geometry collections, regions that only
+ touch the target, and targets without overlap with the source
+- fix averages getting scaled more than once in `tongfen_aggregate` when the metadata lists the same
+ variable name for several datasets
+- fix averages with missing values being pulled toward zero when aggregating with `na.rm = TRUE`,
+ and "Average to" variables sharing a parent variable overwriting each other's base
+- `estimate_tongfen_correspondence` no longer requires the geometry column to be named `geometry`
+- `refresh = TRUE` now also refreshes the cached StatCan correspondence files
+- US Census Bureau relationship files are downloaded to a temporary file before being moved to the cache
+- added missing `\value` documentation for `tongfen_estimate_ca_census` and `tongfen_ca_census_ct`
+- fixed typos in the documentation
# Test environments
* local macOS installation, R 4.6.0
-* GitHub actions (windows-latest, macOS-latest, ubuntu-latest) on release, devel and oldrel
+* GitHub actions: macOS-latest (release), windows-latest (release), ubuntu-latest (devel, release, oldrel-1)
# R CMD check results
0 errors | 0 warnings | 0 notes
diff --git a/data-raw/statcan_correspondence.R b/data-raw/statcan_correspondence.R
new file mode 100644
index 0000000..44e64d5
--- /dev/null
+++ b/data-raw/statcan_correspondence.R
@@ -0,0 +1,51 @@
+## code to prepare the StatCan DA and DB correspondence files hosted on S3
+##
+## StatCan put the correspondence files behind a browser check, so they can't be downloaded
+## programmatically any more. Download and extract the zip files manually from
+## https://www12.statcan.gc.ca/census-recensement/2011/geo/ref/files-fichiers/2006_92-156_DB_ID_txt.zip
+## https://www12.statcan.gc.ca/census-recensement/2011/geo/ref/files-fichiers/2006_92-156_DA_AD_txt.zip
+## https://www12.statcan.gc.ca/census-recensement/2011/geo/ref/files-fichiers/2011_92-156_DB_ID_txt.zip
+## https://www12.statcan.gc.ca/census-recensement/2011/geo/ref/files-fichiers/2011_92-156_DA_AD_txt.zip
+## https://www12.statcan.gc.ca/census-recensement/2011/geo/ref/files-fichiers/2016/2016_92-156_DB_ID_csv.zip
+## https://www12.statcan.gc.ca/census-recensement/2011/geo/ref/files-fichiers/2016/2016_92-156_DA_AD_csv.zip
+## https://www12.statcan.gc.ca/census-recensement/2021/geo/aip-pia/correspondence-correspondance/files-fichiers/2021_92-156-X_DB_ID.zip
+## https://www12.statcan.gc.ca/census-recensement/2021/geo/aip-pia/correspondence-correspondance/files-fichiers/2021_92-156-X_DA_AD.zip
+## into `input_dir`, then run this script and upload the parquet files in `output_dir` to S3.
+
+library(dplyr)
+
+input_dir <- "~/Downloads"
+output_dir <- "~/Downloads/tongfen_statcan_correspondence"
+
+level_codes <- c(DA="DA_AD",DB="DB_ID")
+
+prepare_statcan_correspondence <- function(year,level){
+ new_field <- paste0(level,"UID",year)
+ old_field <- paste0(level,"UID",year-5)
+ dirs <- dir(input_dir,paste0("^",year,"_92-156.*",level_codes[[level]]),full.names=TRUE)
+ dirs <- dirs[dir.exists(dirs)]
+ file <- dir(dirs,"\\.(txt|csv)$",full.names=TRUE,recursive=TRUE)
+ stopifnot(length(file)==1)
+ # DA files are at DB granularity and carry the DBUID in the third column, 2021 files have extra DGUID columns
+ headers <- if (level=="DB") c(new_field,old_field,"flag") else c(new_field,old_field,paste0("DBUID",year),"flag")
+ d <- readr::read_csv(file,col_types=readr::cols(.default="c"),col_names=FALSE) %>%
+ select(all_of(seq_along(headers))) %>%
+ setNames(headers) %>%
+ filter(grepl("^\\d+$",.data[[new_field]])) %>% # header row in the 2016 and 2021 files
+ select(all_of(c(new_field,old_field,"flag"))) %>%
+ unique() %>%
+ arrange(.data[[new_field]],.data[[old_field]])
+ stopifnot(all(grepl("^\\d+$",d[[old_field]])),all(d$flag %in% as.character(1:4)))
+ d
+}
+
+dir.create(output_dir,showWarnings=FALSE)
+for (year in c(2006,2011,2016,2021)) {
+ for (level in c("DA","DB")) {
+ prepare_statcan_correspondence(year,level) %>%
+ nanoparquet::write_parquet(file.path(output_dir,paste0("statcan_correspondence_",year,"_",level,".parquet")),
+ compression="zstd",
+ # nanoparquet does not compress at its default zstd level
+ options=nanoparquet::parquet_options(compression_level=19))
+ }
+}
diff --git a/docs/404.html b/docs/404.html
index d0348fd..008e97f 100644
--- a/docs/404.html
+++ b/docs/404.html
@@ -30,7 +30,7 @@
tongfen
- 0.3.8
+ 0.3.9
As an example, we estimate the share of people in low income in
Vancouver’s skytrain station neighbourhoods. The station neighbourhoods
-are available as part of the cancenus package.
Next we assemble the required metadata for the share of people in
@@ -116,7 +117,7 @@
diff --git a/docs/articles/tongfen-ca-estimate.md b/docs/articles/tongfen-ca-estimate.md
index 2f83abd..6028ceb 100644
--- a/docs/articles/tongfen-ca-estimate.md
+++ b/docs/articles/tongfen-ca-estimate.md
@@ -11,7 +11,7 @@ to estimate census data on custom geographies.
As an example, we estimate the share of people in low income in
Vancouver’s skytrain station neighbourhoods. The station neighbourhoods
-are available as part of the `cancenus` package.
+are available as part of the `cancensus` package.
``` r
diff --git a/docs/articles/tongfen.html b/docs/articles/tongfen.html
index f9f235e..d5e19eb 100644
--- a/docs/articles/tongfen.html
+++ b/docs/articles/tongfen.html
@@ -25,7 +25,7 @@
tongfen
- 0.3.8
+ 0.3.9
@@ -43,6 +43,7 @@
Plotting the cenus tracts for our four census years shows how census
+
Plotting the census tracts for our four census years shows how census
tracts changed over the years.
data%>%
@@ -128,12 +129,12 @@
estimate_tongfen_correspondence function. Unfortunately
this is not an exact science, for example over the years census regions
get adjusted to better align with the road network. Other harmless
-boundary adjustemens can happen along water boundaries, or re-jigging
+boundary adjustments can happen along water boundaries, or re-jigging
boundaries in unpopulated areas.
We are going to impose a tolerance of 200m, where we are calling two
census tract the same if they differ by no more than 200m. We are
specifying that these calculations should be carried out in the
-Statistics Canada Lambert (EPSG:3347) refernce system with units
+Statistics Canada Lambert (EPSG:3347) reference system with units
metres.
diff --git a/docs/articles/tongfen.md b/docs/articles/tongfen.md
index 06301f4..dbc7103 100644
--- a/docs/articles/tongfen.md
+++ b/docs/articles/tongfen.md
@@ -48,7 +48,7 @@ data <- years %>%
}) %>% setNames(years)
```
-Plotting the cenus tracts for our four census years shows how census
+Plotting the census tracts for our four census years shows how census
tracts changed over the years.
``` r
@@ -68,14 +68,15 @@ For this example we will estimate the correspondence between these
regions from the geographic data using the
`estimate_tongfen_correspondence` function. Unfortunately this is not an
exact science, for example over the years census regions get adjusted to
-better align with the road network. Other harmless boundary adjustemens
+better align with the road network. Other harmless boundary adjustments
can happen along water boundaries, or re-jigging boundaries in
unpopulated areas.
We are going to impose a tolerance of 200m, where we are calling two
census tract the same if they differ by no more than 200m. We are
specifying that these calculations should be carried out in the
-Statistics Canada Lambert (EPSG:3347) refernce system with units metres.
+Statistics Canada Lambert (EPSG:3347) reference system with units
+metres.
``` r
@@ -176,7 +177,7 @@ years %>%
It’s time to go back to our original goal of mapping population change.
For this we need to specify how to aggregate up the population data,
which is by simply adding them up. The `meta_for_additive_variables`
-convenience function generates the appropriate metatdata that specifies
+convenience function generates the appropriate metadata that specifies
how to deal with this data.
``` r
diff --git a/docs/articles/tongfen_anomalies.html b/docs/articles/tongfen_anomalies.html
new file mode 100644
index 0000000..a4e6d6f
--- /dev/null
+++ b/docs/articles/tongfen_anomalies.html
@@ -0,0 +1,367 @@
+
+
+
+
+
+
+
+Geocoding anomalies in TongFen timelines • tongfen
+
+
+
+
+
+
+
+
+
+
+
+
+ Skip to contents
+
+
+
TongFen makes data on different geographies comparable by aggregating
+it up to a common geography. The result is only as good as the geocoding
+that assigned the underlying data to geographic regions in the first
+place. Geocoding changes over time, and the same dwelling units, and the
+people living in them, can get assigned to different neighbouring
+regions in different years. In a timeline on a common geography this
+shows up as a surprising drop in one region that is offset by a jump in
+a neighbouring region.
+
The fix is in the spirit of TongFen, joining the affected regions
+gives a slightly coarser geography on which the data is consistent over
+time. The functions in this vignette look for such patterns and join the
+regions on demand. The method is explained in more detail in a blog
+post, this vignette follows the example from that post.
As an example we take the population from the 1971 through 2011
+censuses that Statistics Canada tabulated on 2016 dissemination areas,
+together with the 2016 population. All data comes on the same geography,
+so there is no need to TongFen, but the data for the earlier years is
+geocoded from the road network and block face of the time, which does
+not always match up with 2016 dissemination areas.
Dissemination areas without population in a given year come back as
+missing values. Changes from or to a missing value are never considered
+surprising, so we set them to zero to mark these as areas where nobody
+got counted.
+
The area around Crescent Town shows what the problem looks like.
The population jumps back and forth between the two areas, while the
+sum of the two is fairly steady from 1981 on. People did not move back
+and forth, their homes got geocoded to a different dissemination area in
+different years.
+
+
+
Detecting anomalies
+
+
tongfen_detect_anomalies lists the regions with
+surprising drops. Only decreases are surprising, and a decrease needs to
+be large in both relative and absolute terms. For each of these
+candidate regions it finds the neighbouring region that takes away most
+of the surprise when both are joined, and checks if that reduction is
+large enough to justify joining them. Our regions are identified by
+their GeoUID instead of the TongfenID the
+function looks for by default.
Both areas are candidates, and each one is the neighbour that best
+explains the surprising drops of the other. Not all candidates find a
+neighbour to pair up with. Population does drop for real, for example
+when a site gets cleared for redevelopment, and such regions are left
+alone.
tongfen_anomaly_joins joins the regions that qualify and
+looks again, joined regions can have surprising drops that are
+complemented by another neighbour. This repeats until there are no more
+regions left to join. The result lists the regions that got joined
+together with the identifier of the joined region they are now part of
+and the round in which they first got joined.
tongfen_join_regions applies the joins to the data,
+aggregating the variables and geometries of the regions that get joined
+and leaving all others as they are.
+
+toronto_joined<-tongfen_join_regions(toronto,joins,id="GeoUID")
+
+c(original=nrow(toronto),joined=nrow(toronto_joined))
+#> original joined
+#> 3702 3430
+
+toronto_joined%>%
+filter(GeoUID%in%crescent_town)%>%
+plot_timelines()+
+labs(title="Population in the joined region")
+
+
The map shows the regions that got joined around Crescent Town.
+
+bbox<-toronto%>%filter(GeoUID%in%crescent_town)%>%st_buffer(1500)%>%st_bbox()
+
+ggplot(toronto_joined%>%mutate(joined=GeoUID%in%joins$GeoUID_joined))+
+geom_sf(aes(fill=joined),linewidth=0.1)+
+geom_sf(data=toronto,fill=NA,linewidth=0.1,linetype="dotted")+
+scale_fill_manual(values=c("TRUE"="steelblue","FALSE"="whitesmoke"),guide="none")+
+coord_sf(datum=NA,xlim=bbox[c("xmin","xmax")],ylim=bbox[c("ymin","ymax")])+
+labs(title="Joined regions around Crescent Town",
+ caption="Joined regions in blue, original dissemination areas dotted")
+
+
+
+
Tuning
+
+
Joining regions trades geographic detail for consistency over time,
+and how to best make that trade depends on the data and the application.
+The parameters are documented in tongfen_detect_anomalies,
+the most important ones are
+
+
+rel_scale and abs_scale, the relative and
+absolute decrease at which a change is half way to being fully
+surprising. The defaults of a 25% drop and a drop of 200 are tuned to
+population counts in regions of the size of dissemination areas.
+
+total_surprise_cutoff, how surprising the timeline of a
+region needs to be to become a candidate. The default of 0.75 is
+conservative, above we used 0.4 to also pick up less pronounced
+cases.
+
+cutoff_fact, surprise_reduction_const and
+sum_fact determine how much of the surprise a neighbour
+needs to take away for the regions to get joined.
Neighbours are by default determined by intersecting the geometries
+of the regions. This can miss neighbours if the geometries have been
+simplified, in that case the neighbours argument takes a
+table with the identifiers of neighbouring regions or a neighbours list
+from the spdep package.
+
+
+
Anomalies in TongFen data
+
+
The functions work the same way on data on a common geography built
+by TongFen, where the regions are identified by their
+TongfenID. As an example we look at the dissemination area
+level population in the City of Vancouver for the 2001 through 2021
+censuses.
Passing the metadata to tongfen_join_regions makes sure
+the variables get aggregated the right way, numeric variables that are
+not part of the metadata are assumed to be additive.
Variables that are not additive, like averages, can only be
+aggregated this way if the variable they are averaged over is part of
+the data. The alternative that always works is to join the regions in
+the correspondence the common geography was built from, and use the
+joined correspondence to aggregate the original data. This also is the
+way to use the joins for data other than the one that was used to detect
+the anomalies, for example to get average rents on the corrected
+geography.
+
+
+
+
+
+
+
diff --git a/docs/articles/tongfen_anomalies.md b/docs/articles/tongfen_anomalies.md
new file mode 100644
index 0000000..4e1c12f
--- /dev/null
+++ b/docs/articles/tongfen_anomalies.md
@@ -0,0 +1,311 @@
+# Geocoding anomalies in TongFen timelines
+
+TongFen makes data on different geographies comparable by aggregating it
+up to a common geography. The result is only as good as the geocoding
+that assigned the underlying data to geographic regions in the first
+place. Geocoding changes over time, and the same dwelling units, and the
+people living in them, can get assigned to different neighbouring
+regions in different years. In a timeline on a common geography this
+shows up as a surprising drop in one region that is offset by a jump in
+a neighbouring region.
+
+The fix is in the spirit of TongFen, joining the affected regions gives
+a slightly coarser geography on which the data is consistent over time.
+The functions in this vignette look for such patterns and join the
+regions on demand. The method is explained in more detail in a [blog
+post](https://doodles.mountainmath.ca/posts/2024-07-26-geocoding-errors-in-aggregate-data/),
+this vignette follows the example from that post.
+
+``` r
+
+library(dplyr)
+library(tidyr)
+library(ggplot2)
+library(cancensus)
+library(sf)
+library(tongfen)
+# cancensus::set_api_key("")
+```
+
+## Population timelines for Toronto
+
+As an example we take the population from the 1971 through 2011 censuses
+that Statistics Canada tabulated on 2016 dissemination areas, together
+with the 2016 population. All data comes on the same geography, so there
+is no need to TongFen, but the data for the earlier years is geocoded
+from the road network and block face of the time, which does not always
+match up with 2016 dissemination areas.
+
+``` r
+
+years <- c(1971,seq(1981,2011,5))
+vectors <- c(setNames(paste0("v_CA",years,"x16_1"),years),"2016"="v_CA16_1")
+timeline <- names(vectors)
+
+toronto <- get_census("CA16CT",regions=list(CSD="3520005"),vectors=vectors,
+ level="DA",geo_format="sf",quiet=TRUE) %>%
+ select(GeoUID,all_of(timeline)) %>%
+ mutate(across(all_of(timeline),\(x) coalesce(x,0)))
+```
+
+Dissemination areas without population in a given year come back as
+missing values. Changes from or to a missing value are never considered
+surprising, so we set them to zero to mark these as areas where nobody
+got counted.
+
+The area around Crescent Town shows what the problem looks like.
+
+``` r
+
+crescent_town <- c("35204370","35204765")
+
+plot_timelines <- function(data) {
+ data %>%
+ st_drop_geometry() %>%
+ pivot_longer(all_of(timeline),names_to="Year",values_to="Population") %>%
+ ggplot(aes(x=Year,y=Population,colour=GeoUID,group=GeoUID)) +
+ geom_line() +
+ geom_point() +
+ scale_y_continuous(labels=scales::comma,limits=c(0,NA))
+}
+
+toronto %>%
+ filter(GeoUID %in% crescent_town) %>%
+ plot_timelines() +
+ labs(title="Population in two neighbouring dissemination areas")
+```
+
+
+
+The population jumps back and forth between the two areas, while the sum
+of the two is fairly steady from 1981 on. People did not move back and
+forth, their homes got geocoded to a different dissemination area in
+different years.
+
+## Detecting anomalies
+
+`tongfen_detect_anomalies` lists the regions with surprising drops. Only
+decreases are surprising, and a decrease needs to be large in both
+relative and absolute terms. For each of these candidate regions it
+finds the neighbouring region that takes away most of the surprise when
+both are joined, and checks if that reduction is large enough to justify
+joining them. Our regions are identified by their `GeoUID` instead of
+the `TongfenID` the function looks for by default.
+
+``` r
+
+anomalies <- tongfen_detect_anomalies(toronto,timeline,id="GeoUID",total_surprise_cutoff=0.4)
+
+anomalies %>% filter(GeoUID %in% crescent_town)
+#> # A tibble: 2 × 7
+#> GeoUID surprise_count surprise_total period neighbour surprise_total_joined
+#>
+#> 1 35204370 3 0.933 1991-1… 35204765 0.616
+#> 2 35204765 2 1.03 1981-1… 35204370 0.142
+#> # ℹ 1 more variable: join
+```
+
+Both areas are candidates, and each one is the neighbour that best
+explains the surprising drops of the other. Not all candidates find a
+neighbour to pair up with. Population does drop for real, for example
+when a site gets cleared for redevelopment, and such regions are left
+alone.
+
+``` r
+
+anomalies %>% count(join)
+#> # A tibble: 2 × 2
+#> join n
+#>
+#> 1 FALSE 254
+#> 2 TRUE 288
+```
+
+## Joining regions
+
+`tongfen_anomaly_joins` joins the regions that qualify and looks again,
+joined regions can have surprising drops that are complemented by
+another neighbour. This repeats until there are no more regions left to
+join. The result lists the regions that got joined together with the
+identifier of the joined region they are now part of and the round in
+which they first got joined.
+
+``` r
+
+joins <- tongfen_anomaly_joins(toronto,timeline,id="GeoUID",total_surprise_cutoff=0.4)
+
+joins %>% filter(GeoUID %in% crescent_town)
+#> # A tibble: 2 × 3
+#> GeoUID GeoUID_joined round
+#>
+#> 1 35204370 35204370 1
+#> 2 35204765 35204370 1
+```
+
+`tongfen_join_regions` applies the joins to the data, aggregating the
+variables and geometries of the regions that get joined and leaving all
+others as they are.
+
+``` r
+
+toronto_joined <- tongfen_join_regions(toronto,joins,id="GeoUID")
+
+c(original=nrow(toronto),joined=nrow(toronto_joined))
+#> original joined
+#> 3702 3430
+```
+
+``` r
+
+toronto_joined %>%
+ filter(GeoUID %in% crescent_town) %>%
+ plot_timelines() +
+ labs(title="Population in the joined region")
+```
+
+
+
+The map shows the regions that got joined around Crescent Town.
+
+``` r
+
+bbox <- toronto %>% filter(GeoUID %in% crescent_town) %>% st_buffer(1500) %>% st_bbox()
+
+ggplot(toronto_joined %>% mutate(joined=GeoUID %in% joins$GeoUID_joined)) +
+ geom_sf(aes(fill=joined),linewidth=0.1) +
+ geom_sf(data=toronto,fill=NA,linewidth=0.1,linetype="dotted") +
+ scale_fill_manual(values=c("TRUE"="steelblue","FALSE"="whitesmoke"),guide="none") +
+ coord_sf(datum=NA,xlim=bbox[c("xmin","xmax")],ylim=bbox[c("ymin","ymax")]) +
+ labs(title="Joined regions around Crescent Town",
+ caption="Joined regions in blue, original dissemination areas dotted")
+```
+
+
+
+## Tuning
+
+Joining regions trades geographic detail for consistency over time, and
+how to best make that trade depends on the data and the application. The
+parameters are documented in `tongfen_detect_anomalies`, the most
+important ones are
+
+- `rel_scale` and `abs_scale`, the relative and absolute decrease at
+ which a change is half way to being fully surprising. The defaults of
+ a 25% drop and a drop of 200 are tuned to population counts in regions
+ of the size of dissemination areas.
+- `total_surprise_cutoff`, how surprising the timeline of a region needs
+ to be to become a candidate. The default of 0.75 is conservative,
+ above we used 0.4 to also pick up less pronounced cases.
+- `cutoff_fact`, `surprise_reduction_const` and `sum_fact` determine how
+ much of the surprise a neighbour needs to take away for the regions to
+ get joined.
+
+``` r
+
+c(0.4,0.6,0.75) %>%
+ lapply(\(cutoff) tibble(total_surprise_cutoff=cutoff,
+ regions_joined=tongfen_anomaly_joins(toronto,timeline,id="GeoUID",
+ total_surprise_cutoff=cutoff) %>%
+ nrow())) %>%
+ bind_rows()
+#> # A tibble: 3 × 2
+#> total_surprise_cutoff regions_joined
+#>
+#> 1 0.4 477
+#> 2 0.6 273
+#> 3 0.75 152
+```
+
+Neighbours are by default determined by intersecting the geometries of
+the regions. This can miss neighbours if the geometries have been
+simplified, in that case the `neighbours` argument takes a table with
+the identifiers of neighbouring regions or a neighbours list from the
+**spdep** package.
+
+## Anomalies in TongFen data
+
+The functions work the same way on data on a common geography built by
+TongFen, where the regions are identified by their `TongfenID`. As an
+example we look at the dissemination area level population in the City
+of Vancouver for the 2001 through 2021 censuses.
+
+``` r
+
+regions <- list(CSD="5915022")
+datasets <- c("CA01","CA06","CA11","CA16","CA21")
+meta <- meta_for_additive_variables(datasets,"Population")
+
+vancouver <- get_tongfen_ca_census(regions=regions,meta=meta,level="DA",base_geo="CA21",quiet=TRUE)
+
+joins <- tongfen_anomaly_joins(vancouver,paste0("Population_",datasets),total_surprise_cutoff=0.4)
+joins
+#> # A tibble: 4 × 3
+#> TongfenID TongfenID_joined round
+#>
+#> 1 59150762 59150762 1
+#> 2 59153181 59150762 1
+#> 3 59150765 59150765 1
+#> 4 59150770 59150765 1
+```
+
+Passing the metadata to `tongfen_join_regions` makes sure the variables
+get aggregated the right way, numeric variables that are not part of the
+metadata are assumed to be additive.
+
+``` r
+
+vancouver_joined <- tongfen_join_regions(vancouver,joins,meta)
+```
+
+The `TongfenUID` of the joined regions lists all the dissemination areas
+they are made up of.
+
+``` r
+
+vancouver_joined %>%
+ st_drop_geometry() %>%
+ filter(TongfenID %in% joins$TongfenID_joined) %>%
+ select(TongfenID,TongfenUID,starts_with("Population"))
+#> # A tibble: 2 × 7
+#> TongfenID TongfenUID Population_CA21 Population_CA01 Population_CA06
+#>
+#> 1 59150762 GeoUIDCA01:59150762… 1412 1148 1266
+#> 2 59150765 GeoUIDCA01:59150765… 4082 1895 2499
+#> # ℹ 2 more variables: Population_CA11 , Population_CA16
+```
+
+Variables that are not additive, like averages, can only be aggregated
+this way if the variable they are averaged over is part of the data. The
+alternative that always works is to join the regions in the
+correspondence the common geography was built from, and use the joined
+correspondence to aggregate the original data. This also is the way to
+use the joins for data other than the one that was used to detect the
+anomalies, for example to get average rents on the corrected geography.
+
+``` r
+
+correspondence <- get_tongfen_correspondence_ca_census(geo_datasets=datasets,regions=regions,
+ level="DA",quiet=TRUE) %>%
+ tongfen_join_correspondence(joins)
+
+rent_meta <- meta_for_ca_census_vectors(c(rent_2006="v_CA06_2050",rent_2016="v_CA16_4901"))
+
+rent_data <- c("CA06","CA16") %>%
+ lapply(\(ds) get_census(ds,regions=regions,level="DA",labels="short",quiet=TRUE,
+ vectors=rent_meta %>% filter(geo_dataset==ds) %>% pull(variable),
+ geo_format=if (ds=="CA16") "sf" else NA) %>%
+ rename(!!paste0("GeoUID",ds):="GeoUID")) %>%
+ setNames(c("CA06","CA16"))
+
+rents <- tongfen_aggregate(rent_data,correspondence,rent_meta,base_geo="CA16")
+
+rents %>%
+ st_drop_geometry() %>%
+ filter(TongfenID %in% joins$TongfenID_joined) %>%
+ select(TongfenID,rent_2006,rent_2016)
+#> # A tibble: 2 × 3
+#> TongfenID rent_2006 rent_2016
+#>
+#> 1 59150762 452. 674.
+#> 2 59150765 474. 913.
+```
diff --git a/docs/articles/tongfen_anomalies_files/figure-html/unnamed-chunk-3-1.png b/docs/articles/tongfen_anomalies_files/figure-html/unnamed-chunk-3-1.png
new file mode 100644
index 0000000..f4533c8
Binary files /dev/null and b/docs/articles/tongfen_anomalies_files/figure-html/unnamed-chunk-3-1.png differ
diff --git a/docs/articles/tongfen_anomalies_files/figure-html/unnamed-chunk-8-1.png b/docs/articles/tongfen_anomalies_files/figure-html/unnamed-chunk-8-1.png
new file mode 100644
index 0000000..12a93f2
Binary files /dev/null and b/docs/articles/tongfen_anomalies_files/figure-html/unnamed-chunk-8-1.png differ
diff --git a/docs/articles/tongfen_anomalies_files/figure-html/unnamed-chunk-9-1.png b/docs/articles/tongfen_anomalies_files/figure-html/unnamed-chunk-9-1.png
new file mode 100644
index 0000000..eea4128
Binary files /dev/null and b/docs/articles/tongfen_anomalies_files/figure-html/unnamed-chunk-9-1.png differ
diff --git a/docs/articles/tongfen_ca.html b/docs/articles/tongfen_ca.html
index 8f7bdea..a7a225e 100644
--- a/docs/articles/tongfen_ca.html
+++ b/docs/articles/tongfen_ca.html
@@ -25,7 +25,7 @@
tongfen
- 0.3.8
+ 0.3.9
@@ -43,6 +43,7 @@
The get_tongfen_census_ct and get_tongfen_census_ct_from_da methods make use of the StatCan correspondence files. To speed up this process it is useful to permanently cache these files instead of having to download them repeatedly. If caching is desired, set either
+
The get_tongfen_ca_census and get_tongfen_correspondence_ca_census methods make use of the StatCan correspondence files when run with method = "statcan". Statistics Canada no longer allows programmatic downloads of these files, so the package downloads them from a mirror hosting the files in parquet format. To speed up this process it is useful to permanently cache these files instead of having to download them again in every session. If caching is desired, set either
options("tongfen.cache_path"="<your local cache path>")
Sys.setenv("tongfen.cache_path"="<your local cache path>")
options("custom_data_path"="<your local cache path>")
-
in your .Rprofile or .Renviron file.
+
in your .Rprofile or .Renviron file. Cached files are checked against the mirror once per session and only get downloaded again if they changed, when offline the cached files are used as they are.
The package ships with a subset of voting data from Elections Canada for the 42nd and 43rd federal elections as well as the polling district geographies for the 42nd and 43rd. This facilitates running the example vignette on polling districts without having to download external data. Both are available as open data covered under the Open Government Licence - Canda.
+
The package ships with a subset of voting data from Elections Canada for the 42nd and 43rd federal elections as well as the polling district geographies for the 42nd and 43rd. This facilitates running the example vignette on polling districts without having to download external data. Both are available as open data covered under the Open Government Licence - Canada.
+get_tongfen_us_census integrates the data acquisition (via the tidycensus package) with TongFen, using the US Census Bureau relationship files to build the common geography.
+
+get_tongfen_correspondence_us_census breaks out the correspondence generation from the US Census Bureau relationship files, to tongfen data that comes on census geographies but is obtained by other means. The relationship files are cached in the us_data folder of the cache path described above.
When geographies aren’t sufficiently congruent or the target geography is fixed, we won’t be able to use the tongfen methods to compute the data on a common geography but have to instead rely on estimates. The tongfen_estimate makes no assumption on the underlying geographies and returns estimates of the data on the target geography. It uses area-weighted interpolation to achieve this, and can be refined to dasymmetric estimates using the proportional_reaggregate function.
+
When geographies aren’t sufficiently congruent or the target geography is fixed, we won’t be able to use the tongfen methods to compute the data on a common geography but have to instead rely on estimates. The tongfen_estimate makes no assumption on the underlying geographies and returns estimates of the data on the target geography. It uses area-weighted interpolation to achieve this, and can be refined to dasymetric estimates using the proportional_reaggregate function.
This method has the example that it works independent of the nature of the underlying geographies, but comes at the heavy price of only being an estimate. To be useful for research purposes we also need methods to estimate the errors this introduces and the effects this has on subsequent analysis results.
Methods to facilitate this are still under active development.
@@ -157,14 +160,14 @@
Cite tongfen
If you wish to cite tongfen:
-
von Bergmann, J. (2024). tongfen: R package to Make Data Based on Different Geographies Comparable. v0.3.7. DOI: 10.32614/CRAN.package.tongfen
+
von Bergmann, J. (2026). tongfen: R package to Make Data Based on Different Geographies Comparable. v0.3.9. DOI: 10.32614/CRAN.package.tongfen
A BibTeX entry for LaTeX users is
@Manual{tongfen,
author = {Jens {von Bergmann}},
title = {tongfen: R package to Make Data Based on Different Geographies Comparable},
- year = {2024},
+ year = {2026},
doi = {10.32614/CRAN.package.tongfen},
- note = {R package version 0.3.7},
+ note = {R package version 0.3.9},
url = {https://mountainmath.github.io/tongfen/},
}
@@ -222,7 +225,7 @@
Dev status
diff --git a/docs/index.md b/docs/index.md
index 2d61554..3ffc98b 100644
--- a/docs/index.md
+++ b/docs/index.md
@@ -31,16 +31,21 @@ The latest development version can be installed from GitHub.
### Caching correspondence files
-The `get_tongfen_census_ct` and `get_tongfen_census_ct_from_da` methods
-make use of the StatCan correspondence files. To speed up this process
-it is useful to permanently cache these files instead of having to
-download them repeatedly. If caching is desired, set either
+The `get_tongfen_ca_census` and `get_tongfen_correspondence_ca_census`
+methods make use of the StatCan correspondence files when run with
+`method = "statcan"`. Statistics Canada no longer allows programmatic
+downloads of these files, so the package downloads them from a mirror
+hosting the files in parquet format. To speed up this process it is
+useful to permanently cache these files instead of having to download
+them again in every session. If caching is desired, set either
- `options("tongfen.cache_path"="")`
- `Sys.setenv("tongfen.cache_path"="")`
- `options("custom_data_path"="")`
-in your `.Rprofile` or `.Renviron` file.
+in your `.Rprofile` or `.Renviron` file. Cached files are checked
+against the mirror once per session and only get downloaded again if
+they changed, when offline the cached files are used as they are.
## General TongFen
@@ -90,7 +95,7 @@ and
This facilitates running the example vignette on polling districts
without having to download external data. Both are available as open
data covered under the [Open Government Licence -
-Canda](https://open.canada.ca/en/open-government-licence-canada).
+Canada](https://open.canada.ca/en/open-government-licence-canada).
## Data-specific implementations
@@ -129,8 +134,13 @@ for example [CMHC data](https://www03.cmhc-schl.gc.ca/hmip-pimh).
- `get_tongfen_us_census` integrates the data acquisition (via the
[**tidycensus**
package](https://walker-data.com/tidycensus/index.html)) with TongFen,
- and adds the tongfen `method = "census.gov"` to use the US Census
- Bureau correspondence files for matching.
+ using the US Census Bureau relationship files to build the common
+ geography.
+- `get_tongfen_correspondence_us_census` breaks out the correspondence
+ generation from the US Census Bureau relationship files, to tongfen
+ data that comes on census geographies but is obtained by other means.
+ The relationship files are cached in the `us_data` folder of the cache
+ path described above.
## Other implementations
@@ -145,7 +155,7 @@ data on a common geography but have to instead rely on estimates. The
`tongfen_estimate` makes no assumption on the underlying geographies and
returns estimates of the data on the target geography. It uses
area-weighted interpolation to achieve this, and can be refined to
-dasymmetric estimates using the `proportional_reaggregate` function.
+dasymetric estimates using the `proportional_reaggregate` function.
This method has the example that it works independent of the nature of
the underlying geographies, but comes at the heavy price of only being
@@ -159,8 +169,8 @@ Methods to facilitate this are still under active development.
If you wish to cite tongfen:
-von Bergmann, J. (2024). tongfen: R package to Make Data Based on
-Different Geographies Comparable. v0.3.7. DOI:
+von Bergmann, J. (2026). tongfen: R package to Make Data Based on
+Different Geographies Comparable. v0.3.9. DOI:
10.32614/CRAN.package.tongfen
A BibTeX entry for LaTeX users is
@@ -168,8 +178,8 @@ A BibTeX entry for LaTeX users is
@Manual{tongfen,
author = {Jens {von Bergmann}},
title = {tongfen: R package to Make Data Based on Different Geographies Comparable},
- year = {2024},
+ year = {2026},
doi = {10.32614/CRAN.package.tongfen},
- note = {R package version 0.3.7},
+ note = {R package version 0.3.9},
url = {https://mountainmath.github.io/tongfen/},
}
diff --git a/docs/llms.txt b/docs/llms.txt
index 03a40fe..2f16d0f 100644
--- a/docs/llms.txt
+++ b/docs/llms.txt
@@ -31,16 +31,21 @@ The latest development version can be installed from GitHub.
### Caching correspondence files
-The `get_tongfen_census_ct` and `get_tongfen_census_ct_from_da` methods
-make use of the StatCan correspondence files. To speed up this process
-it is useful to permanently cache these files instead of having to
-download them repeatedly. If caching is desired, set either
+The `get_tongfen_ca_census` and `get_tongfen_correspondence_ca_census`
+methods make use of the StatCan correspondence files when run with
+`method = "statcan"`. Statistics Canada no longer allows programmatic
+downloads of these files, so the package downloads them from a mirror
+hosting the files in parquet format. To speed up this process it is
+useful to permanently cache these files instead of having to download
+them again in every session. If caching is desired, set either
- `options("tongfen.cache_path"="")`
- `Sys.setenv("tongfen.cache_path"="")`
- `options("custom_data_path"="")`
-in your `.Rprofile` or `.Renviron` file.
+in your `.Rprofile` or `.Renviron` file. Cached files are checked
+against the mirror once per session and only get downloaded again if
+they changed, when offline the cached files are used as they are.
## General TongFen
@@ -90,7 +95,7 @@ and
This facilitates running the example vignette on polling districts
without having to download external data. Both are available as open
data covered under the [Open Government Licence -
-Canda](https://open.canada.ca/en/open-government-licence-canada).
+Canada](https://open.canada.ca/en/open-government-licence-canada).
## Data-specific implementations
@@ -129,8 +134,13 @@ for example [CMHC data](https://www03.cmhc-schl.gc.ca/hmip-pimh).
- `get_tongfen_us_census` integrates the data acquisition (via the
[**tidycensus**
package](https://walker-data.com/tidycensus/index.html)) with TongFen,
- and adds the tongfen `method = "census.gov"` to use the US Census
- Bureau correspondence files for matching.
+ using the US Census Bureau relationship files to build the common
+ geography.
+- `get_tongfen_correspondence_us_census` breaks out the correspondence
+ generation from the US Census Bureau relationship files, to tongfen
+ data that comes on census geographies but is obtained by other means.
+ The relationship files are cached in the `us_data` folder of the cache
+ path described above.
## Other implementations
@@ -145,7 +155,7 @@ data on a common geography but have to instead rely on estimates. The
`tongfen_estimate` makes no assumption on the underlying geographies and
returns estimates of the data on the target geography. It uses
area-weighted interpolation to achieve this, and can be refined to
-dasymmetric estimates using the `proportional_reaggregate` function.
+dasymetric estimates using the `proportional_reaggregate` function.
This method has the example that it works independent of the nature of
the underlying geographies, but comes at the heavy price of only being
@@ -159,8 +169,8 @@ Methods to facilitate this are still under active development.
If you wish to cite tongfen:
-von Bergmann, J. (2024). tongfen: R package to Make Data Based on
-Different Geographies Comparable. v0.3.7. DOI:
+von Bergmann, J. (2026). tongfen: R package to Make Data Based on
+Different Geographies Comparable. v0.3.9. DOI:
10.32614/CRAN.package.tongfen
A BibTeX entry for LaTeX users is
@@ -168,9 +178,9 @@ A BibTeX entry for LaTeX users is
@Manual{tongfen,
author = {Jens {von Bergmann}},
title = {tongfen: R package to Make Data Based on Different Geographies Comparable},
- year = {2024},
+ year = {2026},
doi = {10.32614/CRAN.package.tongfen},
- note = {R package version 0.3.7},
+ note = {R package version 0.3.9},
url = {https://mountainmath.github.io/tongfen/},
}
@@ -179,23 +189,23 @@ A BibTeX entry for LaTeX users is
## All functions
- [`add_census_ca_base_variables()`](https://mountainmath.github.io/tongfen/reference/add_census_ca_base_variables.md)
- : Generate metadata from Candian census vectors
+ : Generate metadata from Canadian census vectors
- [`aggregate_data_with_meta()`](https://mountainmath.github.io/tongfen/reference/aggregate_data_with_meta.md)
: Aggregate variables in grouped data
- [`check_tongfen_areas()`](https://mountainmath.github.io/tongfen/reference/check_tongfen_areas.md)
- : Check geographic integrety
+ : Check geographic integrity
- [`check_tongfen_single_areas()`](https://mountainmath.github.io/tongfen/reference/check_tongfen_single_areas.md)
- : Check geographic integrety
+ : Check geographic integrity
- [`estimate_tongfen_correspondence()`](https://mountainmath.github.io/tongfen/reference/estimate_tongfen_correspondence.md)
- : Generate togfen correspondence for list of geographies
+ : Generate tongfen correspondence for list of geographies
- [`estimate_tongfen_single_correspondence()`](https://mountainmath.github.io/tongfen/reference/estimate_tongfen_single_correspondence.md)
- : Generate togfen correspondence for two geographies
+ : Generate tongfen correspondence for two geographies
- [`get_correspondence_ca_census_for()`](https://mountainmath.github.io/tongfen/reference/get_correspondence_ca_census_for.md)
: Get StatCan DA or DB level correspondence file
- [`get_single_correspondence_ca_census_for()`](https://mountainmath.github.io/tongfen/reference/get_single_correspondence_ca_census_for.md)
: Get StatCan DA or DB level correspondence file
- [`get_tongfen_ca_census()`](https://mountainmath.github.io/tongfen/reference/get_tongfen_ca_census.md)
- : Togfen data from several Canadian censuses
+ : Tongfen data from several Canadian censuses
- [`get_tongfen_ca_census_ct_from_da()`](https://mountainmath.github.io/tongfen/reference/get_tongfen_ca_census_ct_from_da.md)
: Canadian census CT level tongfen via DA correspondence
- [`get_tongfen_census_ct()`](https://mountainmath.github.io/tongfen/reference/get_tongfen_census_ct.md)
@@ -207,22 +217,29 @@ A BibTeX entry for LaTeX users is
- [`get_tongfen_correspondence_us_census()`](https://mountainmath.github.io/tongfen/reference/get_tongfen_correspondence_us_census.md)
: Get correspondence table for US census geographies
- [`get_tongfen_us_census()`](https://mountainmath.github.io/tongfen/reference/get_tongfen_us_census.md)
- : Get US census data for 2000 and 2010 census on common census tract
- based geography
+ : Get US census data for several censuses on a common geography
- [`meta_for_additive_variables()`](https://mountainmath.github.io/tongfen/reference/meta_for_additive_variables.md)
: Generate tongfen metadata for additive variables
- [`meta_for_ca_census_vectors()`](https://mountainmath.github.io/tongfen/reference/meta_for_ca_census_vectors.md)
- : Generate metadata from Candian census vectors
+ : Generate metadata from Canadian census vectors
- [`proportional_reaggregate()`](https://mountainmath.github.io/tongfen/reference/proportional_reaggregate.md)
: Dasymetric downsampling
- [`tongfen_aggregate()`](https://mountainmath.github.io/tongfen/reference/tongfen_aggregate.md)
: Perform tongfen according to correspondence
+- [`tongfen_anomaly_joins()`](https://mountainmath.github.io/tongfen/reference/tongfen_anomaly_joins.md)
+ : Determine regions to join to correct for likely geocoding anomalies
- [`tongfen_ca_census_ct()`](https://mountainmath.github.io/tongfen/reference/tongfen_ca_census_ct.md)
: Canadian census CT level tongfen via identifier matching
+- [`tongfen_detect_anomalies()`](https://mountainmath.github.io/tongfen/reference/tongfen_detect_anomalies.md)
+ : Detect likely geocoding anomalies in timelines on a common geography
- [`tongfen_estimate()`](https://mountainmath.github.io/tongfen/reference/tongfen_estimate.md)
: Estimate variable values for custom geography
- [`tongfen_estimate_ca_census()`](https://mountainmath.github.io/tongfen/reference/tongfen_estimate_ca_census.md)
: Tongfen estimate data for given geometry
+- [`tongfen_join_correspondence()`](https://mountainmath.github.io/tongfen/reference/tongfen_join_correspondence.md)
+ : Join regions in a correspondence
+- [`tongfen_join_regions()`](https://mountainmath.github.io/tongfen/reference/tongfen_join_regions.md)
+ : Join regions in data on a common geography
- [`tongfen_tag_largest_overlap()`](https://mountainmath.github.io/tongfen/reference/tongfen_tag_largest_overlap.md)
: Tag regions by largest overlap
- [`vancouver_elections_data_2015`](https://mountainmath.github.io/tongfen/reference/vancouver_elections_data_2015.md)
@@ -261,3 +278,8 @@ A BibTeX entry for LaTeX users is
- [TongFen for US census
data](https://mountainmath.github.io/tongfen/articles/tongfen_us.md):
+
+### Geocoding anomalies
+
+- [Geocoding anomalies in TongFen
+ timelines](https://mountainmath.github.io/tongfen/articles/tongfen_anomalies.md):
diff --git a/docs/news/index.html b/docs/news/index.html
index ecc6397..ba0514c 100644
--- a/docs/news/index.html
+++ b/docs/news/index.html
@@ -7,7 +7,7 @@
tongfen
- 0.3.8
+ 0.3.9
@@ -23,6 +23,7 @@
-Joins the StatCan correspodence files for several census years
+Joins the StatCan correspondence files for several census years
@@ -85,7 +86,7 @@
Value
diff --git a/docs/reference/get_correspondence_ca_census_for.md b/docs/reference/get_correspondence_ca_census_for.md
index 7d49eb6..326a8d1 100644
--- a/docs/reference/get_correspondence_ca_census_for.md
+++ b/docs/reference/get_correspondence_ca_census_for.md
@@ -1,6 +1,6 @@
# Get StatCan DA or DB level correspondence file
-**\[deprecated\]** Joins the StatCan correspodence files for several
+**\[deprecated\]** Joins the StatCan correspondence files for several
census years
## Usage
diff --git a/docs/reference/get_single_correspondence_ca_census_for.html b/docs/reference/get_single_correspondence_ca_census_for.html
index 09bf3a9..0f93e8c 100644
--- a/docs/reference/get_single_correspondence_ca_census_for.html
+++ b/docs/reference/get_single_correspondence_ca_census_for.html
@@ -1,5 +1,13 @@
-Get StatCan DA or DB level correspondence file — get_single_correspondence_ca_census_for • tongfen
+Get StatCan DA or DB level correspondence file — get_single_correspondence_ca_census_for • tongfenSkip to contents
@@ -7,7 +15,7 @@
tongfen
- 0.3.8
+ 0.3.9
@@ -23,6 +31,7 @@
The correspondence files are downloaded from a mirror of the Statistics Canada correspondence files
+and cached in the tongfen cache directory. The cached files are checked against the mirror once per session
+and get downloaded again if they changed. The location of the mirror can be changed via the
+`tongfen.statcan_correspondence_url` option.
@@ -83,7 +96,7 @@
Value
diff --git a/docs/reference/get_single_correspondence_ca_census_for.md b/docs/reference/get_single_correspondence_ca_census_for.md
index 89b9fe6..331d1ee 100644
--- a/docs/reference/get_single_correspondence_ca_census_for.md
+++ b/docs/reference/get_single_correspondence_ca_census_for.md
@@ -2,6 +2,12 @@
**\[maturing\]**
+The correspondence files are downloaded from a mirror of the Statistics
+Canada correspondence files and cached in the tongfen cache directory.
+The cached files are checked against the mirror once per session and get
+downloaded again if they changed. The location of the mirror can be
+changed via the \`tongfen.statcan_correspondence_url\` option.
+
## Usage
``` r
diff --git a/docs/reference/get_tongfen_ca_census.html b/docs/reference/get_tongfen_ca_census.html
index 498a3f5..f257651 100644
--- a/docs/reference/get_tongfen_ca_census.html
+++ b/docs/reference/get_tongfen_ca_census.html
@@ -1,7 +1,7 @@
-Togfen data from several Canadian censuses — get_tongfen_ca_census • tongfen
+Tongfen data from several Canadian censuses — get_tongfen_ca_census • tongfenSkip to contents
@@ -9,7 +9,7 @@
tongfen
- 0.3.8
+ 0.3.9
@@ -25,6 +25,7 @@
diff --git a/docs/reference/get_tongfen_ca_census.md b/docs/reference/get_tongfen_ca_census.md
index fbf9e44..daeb89f 100644
--- a/docs/reference/get_tongfen_ca_census.md
+++ b/docs/reference/get_tongfen_ca_census.md
@@ -1,8 +1,8 @@
-# Togfen data from several Canadian censuses
+# Tongfen data from several Canadian censuses
**\[maturing\]**
-Get data from several Candian censuses on a common geography. Requires
+Get data from several Canadian censuses on a common geography. Requires
sf and cancensus package to be available
## Usage
@@ -32,7 +32,7 @@ get_tongfen_ca_census(
- meta:
- metadata for the census veraiables to aggregate, for example as
+ metadata for the census variables to aggregate, for example as
returned by `meta_for_ca_census_vectors`.
- level:
@@ -62,7 +62,7 @@ get_tongfen_ca_census(
- tolerance:
- tolerance for \`estimate_tongen_correspondence\` in metres, default
+ tolerance for \`estimate_tongfen_correspondence\` in metres, default
value is 50 metres, only used when method is 'estimate' or
'identifier'
diff --git a/docs/reference/get_tongfen_ca_census_ct_from_da.html b/docs/reference/get_tongfen_ca_census_ct_from_da.html
index 68370f3..23092e2 100644
--- a/docs/reference/get_tongfen_ca_census_ct_from_da.html
+++ b/docs/reference/get_tongfen_ca_census_ct_from_da.html
@@ -11,7 +11,7 @@
tongfen
- 0.3.8
+ 0.3.9
@@ -27,6 +27,7 @@
diff --git a/docs/reference/get_tongfen_correspondence_ca_census.html b/docs/reference/get_tongfen_correspondence_ca_census.html
index a0447f7..6ba7914 100644
--- a/docs/reference/get_tongfen_correspondence_ca_census.html
+++ b/docs/reference/get_tongfen_correspondence_ca_census.html
@@ -1,7 +1,7 @@
Get StatCan correspondence data — get_tongfen_correspondence_ca_census • tongfen
+Get correspondence file for several Canadian censuses on a common geography. Requires sf and cancensus package to be available">
Skip to contents
@@ -9,7 +9,7 @@
tongfen
- 0.3.8
+ 0.3.9
@@ -25,6 +25,7 @@
diff --git a/docs/reference/get_tongfen_correspondence_ca_census.md b/docs/reference/get_tongfen_correspondence_ca_census.md
index 9a6bafc..aee34af 100644
--- a/docs/reference/get_tongfen_correspondence_ca_census.md
+++ b/docs/reference/get_tongfen_correspondence_ca_census.md
@@ -2,7 +2,7 @@
**\[maturing\]**
-Get correspondence file for several Candian censuses on a common
+Get correspondence file for several Canadian censuses on a common
geography. Requires sf and cancensus package to be available
## Usage
@@ -48,7 +48,7 @@ get_tongfen_correspondence_ca_census(
- tolerance:
- tolerance for \`estimate_tongen_correspondence\` in metres, default
+ tolerance for \`estimate_tongfen_correspondence\` in metres, default
value is 50 metres, only used when method is 'estimate' or
'identifier'
diff --git a/docs/reference/get_tongfen_correspondence_us_census.html b/docs/reference/get_tongfen_correspondence_us_census.html
index 5b80c5b..bc78b58 100644
--- a/docs/reference/get_tongfen_correspondence_us_census.html
+++ b/docs/reference/get_tongfen_correspondence_us_census.html
@@ -31,7 +31,7 @@
tongfen
- 0.3.8
+ 0.3.9
@@ -47,6 +47,7 @@
optional path to cache the relationship files in, defaults to the
-`tongfen.cache_path` option and falls back to a temporary directory
+`tongfen.cache_path` option. If that is not set the `tongfen.cache_path` environment variable
+and the `custom_data_path` option are used, falling back to a temporary directory
diff --git a/docs/reference/get_tongfen_correspondence_us_census.md b/docs/reference/get_tongfen_correspondence_us_census.md
index aeaeede..671a3cf 100644
--- a/docs/reference/get_tongfen_correspondence_us_census.md
+++ b/docs/reference/get_tongfen_correspondence_us_census.md
@@ -67,7 +67,10 @@ get_tongfen_correspondence_us_census(
- cache_path:
optional path to cache the relationship files in, defaults to the
- \`tongfen.cache_path\` option and falls back to a temporary directory
+ \`tongfen.cache_path\` option. If that is not set the
+ \`tongfen.cache_path\` environment variable and the
+ \`custom_data_path\` option are used, falling back to a temporary
+ directory
## Value
diff --git a/docs/reference/get_tongfen_us_census.html b/docs/reference/get_tongfen_us_census.html
index 57f67e0..1dcc3e1 100644
--- a/docs/reference/get_tongfen_us_census.html
+++ b/docs/reference/get_tongfen_us_census.html
@@ -1,5 +1,5 @@
-Get US census data for 2000 and 2010 census on common census tract based geography — get_tongfen_us_census • tongfenGet US census data for several censuses on a common geography — get_tongfen_us_census • tongfentongfen
- 0.3.8
+ 0.3.9
@@ -35,6 +35,7 @@
census year to use as base geography, default is `2010`.
+
dataset to use as base geography, for example `"dec2010"`, has to be one of
+the datasets in `meta`. Default is `NULL`, which uses the first dataset in `meta`.
diff --git a/docs/reference/get_tongfen_us_census.md b/docs/reference/get_tongfen_us_census.md
index 8d5e539..3b22fd0 100644
--- a/docs/reference/get_tongfen_us_census.md
+++ b/docs/reference/get_tongfen_us_census.md
@@ -1,4 +1,4 @@
-# Get US census data for 2000 and 2010 census on common census tract based geography
+# Get US census data for several censuses on a common geography
**\[maturing\]**
@@ -48,7 +48,9 @@ get_tongfen_us_census(
- base_geo:
- census year to use as base geography, default is \`2010\`.
+ dataset to use as base geography, for example \`"dec2010"\`, has to be
+ one of the datasets in \`meta\`. Default is \`NULL\`, which uses the
+ first dataset in \`meta\`.
- min_area_share:
@@ -68,7 +70,7 @@ get_tongfen_us_census(
## Value
sf object with (wide form) census variables with census year as suffix
-(separated by underdcore "\_").
+(separated by underscore "\_").
## Examples
diff --git a/docs/reference/index.html b/docs/reference/index.html
index 0ea61c5..69c3312 100644
--- a/docs/reference/index.html
+++ b/docs/reference/index.html
@@ -7,7 +7,7 @@
tongfen
- 0.3.8
+ 0.3.9
@@ -23,6 +23,7 @@
diff --git a/docs/reference/index.md b/docs/reference/index.md
index 6be6a15..2fd4b3a 100644
--- a/docs/reference/index.md
+++ b/docs/reference/index.md
@@ -3,23 +3,23 @@
## All functions
- [`add_census_ca_base_variables()`](https://mountainmath.github.io/tongfen/reference/add_census_ca_base_variables.md)
- : Generate metadata from Candian census vectors
+ : Generate metadata from Canadian census vectors
- [`aggregate_data_with_meta()`](https://mountainmath.github.io/tongfen/reference/aggregate_data_with_meta.md)
: Aggregate variables in grouped data
- [`check_tongfen_areas()`](https://mountainmath.github.io/tongfen/reference/check_tongfen_areas.md)
- : Check geographic integrety
+ : Check geographic integrity
- [`check_tongfen_single_areas()`](https://mountainmath.github.io/tongfen/reference/check_tongfen_single_areas.md)
- : Check geographic integrety
+ : Check geographic integrity
- [`estimate_tongfen_correspondence()`](https://mountainmath.github.io/tongfen/reference/estimate_tongfen_correspondence.md)
- : Generate togfen correspondence for list of geographies
+ : Generate tongfen correspondence for list of geographies
- [`estimate_tongfen_single_correspondence()`](https://mountainmath.github.io/tongfen/reference/estimate_tongfen_single_correspondence.md)
- : Generate togfen correspondence for two geographies
+ : Generate tongfen correspondence for two geographies
- [`get_correspondence_ca_census_for()`](https://mountainmath.github.io/tongfen/reference/get_correspondence_ca_census_for.md)
: Get StatCan DA or DB level correspondence file
- [`get_single_correspondence_ca_census_for()`](https://mountainmath.github.io/tongfen/reference/get_single_correspondence_ca_census_for.md)
: Get StatCan DA or DB level correspondence file
- [`get_tongfen_ca_census()`](https://mountainmath.github.io/tongfen/reference/get_tongfen_ca_census.md)
- : Togfen data from several Canadian censuses
+ : Tongfen data from several Canadian censuses
- [`get_tongfen_ca_census_ct_from_da()`](https://mountainmath.github.io/tongfen/reference/get_tongfen_ca_census_ct_from_da.md)
: Canadian census CT level tongfen via DA correspondence
- [`get_tongfen_census_ct()`](https://mountainmath.github.io/tongfen/reference/get_tongfen_census_ct.md)
@@ -31,22 +31,29 @@
- [`get_tongfen_correspondence_us_census()`](https://mountainmath.github.io/tongfen/reference/get_tongfen_correspondence_us_census.md)
: Get correspondence table for US census geographies
- [`get_tongfen_us_census()`](https://mountainmath.github.io/tongfen/reference/get_tongfen_us_census.md)
- : Get US census data for 2000 and 2010 census on common census tract
- based geography
+ : Get US census data for several censuses on a common geography
- [`meta_for_additive_variables()`](https://mountainmath.github.io/tongfen/reference/meta_for_additive_variables.md)
: Generate tongfen metadata for additive variables
- [`meta_for_ca_census_vectors()`](https://mountainmath.github.io/tongfen/reference/meta_for_ca_census_vectors.md)
- : Generate metadata from Candian census vectors
+ : Generate metadata from Canadian census vectors
- [`proportional_reaggregate()`](https://mountainmath.github.io/tongfen/reference/proportional_reaggregate.md)
: Dasymetric downsampling
- [`tongfen_aggregate()`](https://mountainmath.github.io/tongfen/reference/tongfen_aggregate.md)
: Perform tongfen according to correspondence
+- [`tongfen_anomaly_joins()`](https://mountainmath.github.io/tongfen/reference/tongfen_anomaly_joins.md)
+ : Determine regions to join to correct for likely geocoding anomalies
- [`tongfen_ca_census_ct()`](https://mountainmath.github.io/tongfen/reference/tongfen_ca_census_ct.md)
: Canadian census CT level tongfen via identifier matching
+- [`tongfen_detect_anomalies()`](https://mountainmath.github.io/tongfen/reference/tongfen_detect_anomalies.md)
+ : Detect likely geocoding anomalies in timelines on a common geography
- [`tongfen_estimate()`](https://mountainmath.github.io/tongfen/reference/tongfen_estimate.md)
: Estimate variable values for custom geography
- [`tongfen_estimate_ca_census()`](https://mountainmath.github.io/tongfen/reference/tongfen_estimate_ca_census.md)
: Tongfen estimate data for given geometry
+- [`tongfen_join_correspondence()`](https://mountainmath.github.io/tongfen/reference/tongfen_join_correspondence.md)
+ : Join regions in a correspondence
+- [`tongfen_join_regions()`](https://mountainmath.github.io/tongfen/reference/tongfen_join_regions.md)
+ : Join regions in data on a common geography
- [`tongfen_tag_largest_overlap()`](https://mountainmath.github.io/tongfen/reference/tongfen_tag_largest_overlap.md)
: Tag regions by largest overlap
- [`vancouver_elections_data_2015`](https://mountainmath.github.io/tongfen/reference/vancouver_elections_data_2015.md)
diff --git a/docs/reference/meta_for_additive_variables.html b/docs/reference/meta_for_additive_variables.html
index 7c0269a..f4a2a69 100644
--- a/docs/reference/meta_for_additive_variables.html
+++ b/docs/reference/meta_for_additive_variables.html
@@ -9,7 +9,7 @@
tongfen
- 0.3.8
+ 0.3.9
@@ -25,6 +25,7 @@
Column name to use for proportional weighting when re-aggregating, or named vector with column name for each category.
-Categries that should be re-aggregated as means should be set to NA and will only be reaggregated if the base data has NA values.
+Categories that should be re-aggregated as means should be set to NA and will only be reaggregated if the base data has NA values.
diff --git a/docs/reference/proportional_reaggregate.md b/docs/reference/proportional_reaggregate.md
index d2882f4..91be6da 100644
--- a/docs/reference/proportional_reaggregate.md
+++ b/docs/reference/proportional_reaggregate.md
@@ -46,8 +46,8 @@ proportional_reaggregate(
- base:
Column name to use for proportional weighting when re-aggregating, or
- named vector with column name for each category. Categries that should
- be re-aggregated as means should be set to NA and will only be
+ named vector with column name for each category. Categories that
+ should be re-aggregated as means should be set to NA and will only be
reaggregated if the base data has NA values.
## Value
diff --git a/docs/reference/tongfen_aggregate.html b/docs/reference/tongfen_aggregate.html
index 0f3d2ae..2f6dde8 100644
--- a/docs/reference/tongfen_aggregate.html
+++ b/docs/reference/tongfen_aggregate.html
@@ -1,7 +1,7 @@
Perform tongfen according to correspondence — tongfen_aggregate • tongfen
+Aggregate variables specified in meta for several datasets according to correspondence.">
Skip to contents
@@ -9,7 +9,7 @@
tongfen
- 0.3.8
+ 0.3.9
@@ -25,6 +25,7 @@
named list of datasets to be aggregated. The names identify the datasets, they are
+matched against the `geo_dataset` column in `meta` to pick the aggregation rules and labels
+for each dataset. Without names, or with names not found in `meta`, the rules for all
+datasets are applied and the variables keep their original names
correspondence
@@ -92,17 +96,17 @@
Value
Examples
-
# aggregate census tract level 2006 population data on common gepgraphy build through
+
# aggregate census tract level 2006 and 2016 population data on common geography built through# correspondence from 2006 and 2016 census tracts in the City of Vancouver.if(FALSE){# \dontrun{regions<-list(CSD="5915022")geo1<-cancensus::get_census("CA06",regions=regions,geo_format='sf',level='CT')geo2<-cancensus::get_census("CA16",regions=regions,geo_format='sf',level='CT')
-meta<-meta_for_additive_variables("CA06","Population")
+meta<-meta_for_additive_variables(c("CA06","CA16"),"Population")correspondence<-get_tongfen_correspondence_ca_census(geo_datasets=c('CA06','CA16'), regions=regions,level='CT')
-result<-tongfen_aggregate(list(geo1%>%rename(GeoUIDCA06=GeoUID),
-geo2%>%rename(GeoUIDCA16=GeoUID)),correspondence,meta)
+result<-tongfen_aggregate(list(CA06=geo1%>%rename(GeoUIDCA06=GeoUID),
+ CA16=geo2%>%rename(GeoUIDCA16=GeoUID)),correspondence,meta)}# }
diff --git a/docs/reference/tongfen_aggregate.md b/docs/reference/tongfen_aggregate.md
index 762d736..3853029 100644
--- a/docs/reference/tongfen_aggregate.md
+++ b/docs/reference/tongfen_aggregate.md
@@ -2,7 +2,7 @@
**\[maturing\]**
-Aggregate variables secified in meta for several datasets according to
+Aggregate variables specified in meta for several datasets according to
correspondence.
## Usage
@@ -21,7 +21,11 @@ tongfen_aggregate(
- data:
- list of datasets to be aggregated
+ named list of datasets to be aggregated. The names identify the
+ datasets, they are matched against the \`geo_dataset\` column in
+ \`meta\` to pick the aggregation rules and labels for each dataset.
+ Without names, or with names not found in \`meta\`, the rules for all
+ datasets are applied and the variables keep their original names
- correspondence:
@@ -51,16 +55,16 @@ type sf or tibble otherwise.
## Examples
``` r
-# aggregate census tract level 2006 population data on common gepgraphy build through
+# aggregate census tract level 2006 and 2016 population data on common geography built through
# correspondence from 2006 and 2016 census tracts in the City of Vancouver.
if (FALSE) { # \dontrun{
regions <- list(CSD="5915022")
geo1 <- cancensus::get_census("CA06",regions=regions,geo_format='sf',level='CT')
geo2 <- cancensus::get_census("CA16",regions=regions,geo_format='sf',level='CT')
-meta <- meta_for_additive_variables("CA06","Population")
+meta <- meta_for_additive_variables(c("CA06","CA16"),"Population")
correspondence <- get_tongfen_correspondence_ca_census(geo_datasets=c('CA06','CA16'),
regions=regions,level='CT')
-result <- tongfen_aggregate(list(geo1 %>% rename(GeoUIDCA06=GeoUID),
- geo2 %>% rename(GeoUIDCA16=GeoUID)),correspondence,meta)
+result <- tongfen_aggregate(list(CA06=geo1 %>% rename(GeoUIDCA06=GeoUID),
+ CA16=geo2 %>% rename(GeoUIDCA16=GeoUID)),correspondence,meta)
} # }
```
diff --git a/docs/reference/tongfen_anomaly_joins.html b/docs/reference/tongfen_anomaly_joins.html
new file mode 100644
index 0000000..6f8456b
--- /dev/null
+++ b/docs/reference/tongfen_anomaly_joins.html
@@ -0,0 +1,202 @@
+
+Determine regions to join to correct for likely geocoding anomalies — tongfen_anomaly_joins • tongfen
+ Skip to contents
+
+
+
+
+
+
Determine regions to join to correct for likely geocoding anomalies
Looks for regions with surprising drops in the timeline of a count variable that are complemented by
+a neighbouring region, as explained in `tongfen_detect_anomalies`, and joins them. This gets
+repeated on the joined regions until there are no more regions left that qualify to get joined.
+In each round a region only gets joined with one other region, the most surprising regions go first.
+
Joining regions trades geographic detail for timelines that are consistent over time. The parameters
+control how aggressively regions get joined and are best calibrated on the data at hand, erring on the side of
+joining too few regions risks keeping geocoding problems, erring on the other side risks removing real
+changes and needlessly coarsens the geography.
+
The result can be used to join the regions via `tongfen_join_regions`, or to update a correspondence
+via `tongfen_join_correspondence`.
data on a common geography, with one row per region, for example as returned by
+`tongfen_aggregate` or `get_tongfen_ca_census`. Needs to be of class sf unless `neighbours` is specified.
+
+
+
variables
+
names of the columns holding the timeline of a count variable like population or
+dwellings, in temporal order. Changes from or to a missing value are not surprising and don't make up for
+surprising changes in neighbouring regions, and joined regions are missing a value if one of the regions they
+are made up of is. Replace missing values by zero beforehand if they stand for regions where nothing got counted
+
+
+
id
+
name of the column that uniquely identifies the regions, default is "TongfenID"
+
+
+
neighbours
+
optional, neighbouring regions as a table with the identifiers of pairs of neighbouring
+regions in the first two columns, or as a neighbours list like the ones returned by `spdep::poly2nb`.
+By default all regions with intersecting geometries are neighbours, which can miss neighbours
+if the geometries have been simplified and don't share their boundaries any more.
+
+
+
rel_scale
+
relative decrease that is half way to full surprise, default is `0.25` for a 25% drop
+
+
+
abs_scale
+
absolute decrease that is half way to full surprise, default is `200`
+
+
+
p
+
exponent of the norm used to combine the surprises across the timeline into the total surprise.
+Large values focus on the most surprising change, 1 adds up the surprises of all changes, default is `4`
+
+
+
surprise_cutoff
+
changes with larger surprise count as surprising, default is `0.15`. Only
+regions with at least one surprising change are candidates
+
+
+
total_surprise_cutoff
+
only regions with larger total surprise are candidates, default is `0.75`
+
+
+
cutoff_fact
+
join regions if the total surprise after joining is lower than this share of the
+total surprise of the candidate region, default is `0.6`
+
+
+
surprise_reduction_const
+
join regions if joining lowers the total surprise by more than this,
+default is `0.15`
+
+
+
sum_fact
+
join regions if the total surprise after joining is lower than `cutoff_fact * sum_fact`
+times the sum of the total surprises of both regions, default is `0.7`
+
+
+
+
Value
+
A tibble with one row for each region that gets joined with other regions, with the identifier
+of the region, the identifier of the joined region it becomes part of in the column named like the
+identifier with suffix `_joined`, by default `TongfenID_joined`, and the `round` in which the region
+first got joined to another region. The identifier of a joined region is the smallest
+identifier of the regions it is made up of.
+
+
+
+
Examples
+
# Correct 2001 through 2021 dissemination area level population timelines in the
+# City of Vancouver for likely geocoding problems
+if(FALSE){# \dontrun{
+datasets<-c("CA01","CA06","CA11","CA16","CA21")
+meta<-meta_for_additive_variables(datasets,"Population")
+data<-get_tongfen_ca_census(regions=list(CSD="5915022"),meta=meta,level="DA",base_geo="CA21")
+
+joins<-tongfen_anomaly_joins(data,paste0("Population_",datasets))
+corrected_data<-tongfen_join_regions(data,joins,meta)
+}# }
+
+
+
+
+
+
+
+
+
+
+
+
+
diff --git a/docs/reference/tongfen_anomaly_joins.md b/docs/reference/tongfen_anomaly_joins.md
new file mode 100644
index 0000000..5e39412
--- /dev/null
+++ b/docs/reference/tongfen_anomaly_joins.md
@@ -0,0 +1,139 @@
+# Determine regions to join to correct for likely geocoding anomalies
+
+**\[experimental\]**
+
+Looks for regions with surprising drops in the timeline of a count
+variable that are complemented by a neighbouring region, as explained in
+\`tongfen_detect_anomalies\`, and joins them. This gets repeated on the
+joined regions until there are no more regions left that qualify to get
+joined. In each round a region only gets joined with one other region,
+the most surprising regions go first.
+
+Joining regions trades geographic detail for timelines that are
+consistent over time. The parameters control how aggressively regions
+get joined and are best calibrated on the data at hand, erring on the
+side of joining too few regions risks keeping geocoding problems, erring
+on the other side risks removing real changes and needlessly coarsens
+the geography.
+
+The result can be used to join the regions via \`tongfen_join_regions\`,
+or to update a correspondence via \`tongfen_join_correspondence\`.
+
+## Usage
+
+``` r
+tongfen_anomaly_joins(
+ data,
+ variables,
+ id = "TongfenID",
+ neighbours = NULL,
+ rel_scale = 0.25,
+ abs_scale = 200,
+ p = 4,
+ surprise_cutoff = 0.15,
+ total_surprise_cutoff = 0.75,
+ cutoff_fact = 0.6,
+ surprise_reduction_const = 0.15,
+ sum_fact = 0.7
+)
+```
+
+## Arguments
+
+- data:
+
+ data on a common geography, with one row per region, for example as
+ returned by \`tongfen_aggregate\` or \`get_tongfen_ca_census\`. Needs
+ to be of class sf unless \`neighbours\` is specified.
+
+- variables:
+
+ names of the columns holding the timeline of a count variable like
+ population or dwellings, in temporal order. Changes from or to a
+ missing value are not surprising and don't make up for surprising
+ changes in neighbouring regions, and joined regions are missing a
+ value if one of the regions they are made up of is. Replace missing
+ values by zero beforehand if they stand for regions where nothing got
+ counted
+
+- id:
+
+ name of the column that uniquely identifies the regions, default is
+ "TongfenID"
+
+- neighbours:
+
+ optional, neighbouring regions as a table with the identifiers of
+ pairs of neighbouring regions in the first two columns, or as a
+ neighbours list like the ones returned by \`spdep::poly2nb\`. By
+ default all regions with intersecting geometries are neighbours, which
+ can miss neighbours if the geometries have been simplified and don't
+ share their boundaries any more.
+
+- rel_scale:
+
+ relative decrease that is half way to full surprise, default is
+ \`0.25\` for a 25% drop
+
+- abs_scale:
+
+ absolute decrease that is half way to full surprise, default is
+ \`200\`
+
+- p:
+
+ exponent of the norm used to combine the surprises across the timeline
+ into the total surprise. Large values focus on the most surprising
+ change, 1 adds up the surprises of all changes, default is \`4\`
+
+- surprise_cutoff:
+
+ changes with larger surprise count as surprising, default is \`0.15\`.
+ Only regions with at least one surprising change are candidates
+
+- total_surprise_cutoff:
+
+ only regions with larger total surprise are candidates, default is
+ \`0.75\`
+
+- cutoff_fact:
+
+ join regions if the total surprise after joining is lower than this
+ share of the total surprise of the candidate region, default is
+ \`0.6\`
+
+- surprise_reduction_const:
+
+ join regions if joining lowers the total surprise by more than this,
+ default is \`0.15\`
+
+- sum_fact:
+
+ join regions if the total surprise after joining is lower than
+ \`cutoff_fact \* sum_fact\` times the sum of the total surprises of
+ both regions, default is \`0.7\`
+
+## Value
+
+A tibble with one row for each region that gets joined with other
+regions, with the identifier of the region, the identifier of the joined
+region it becomes part of in the column named like the identifier with
+suffix \`\_joined\`, by default \`TongfenID_joined\`, and the \`round\`
+in which the region first got joined to another region. The identifier
+of a joined region is the smallest identifier of the regions it is made
+up of.
+
+## Examples
+
+``` r
+# Correct 2001 through 2021 dissemination area level population timelines in the
+# City of Vancouver for likely geocoding problems
+if (FALSE) { # \dontrun{
+datasets <- c("CA01","CA06","CA11","CA16","CA21")
+meta <- meta_for_additive_variables(datasets,"Population")
+data <- get_tongfen_ca_census(regions=list(CSD="5915022"),meta=meta,level="DA",base_geo="CA21")
+
+joins <- tongfen_anomaly_joins(data,paste0("Population_",datasets))
+corrected_data <- tongfen_join_regions(data,joins,meta)
+} # }
+```
diff --git a/docs/reference/tongfen_ca_census_ct.html b/docs/reference/tongfen_ca_census_ct.html
index cc9ebcc..090d3f2 100644
--- a/docs/reference/tongfen_ca_census_ct.html
+++ b/docs/reference/tongfen_ca_census_ct.html
@@ -9,7 +9,7 @@
tongfen
- 0.3.8
+ 0.3.9
@@ -25,6 +25,7 @@
diff --git a/docs/reference/tongfen_ca_census_ct.md b/docs/reference/tongfen_ca_census_ct.md
index 4254073..4b4ccbf 100644
--- a/docs/reference/tongfen_ca_census_ct.md
+++ b/docs/reference/tongfen_ca_census_ct.md
@@ -21,12 +21,12 @@ tongfen_ca_census_ct(
- data1:
- cancensus CT level datatset for year1 \< year2 to serve as base for
+ cancensus CT level dataset for year1 \< year2 to serve as base for
common geography
- data2:
- cancensus CT level datatset for year2 to be aggregated to common
+ cancensus CT level dataset for year2 to be aggregated to common
geography
- data2_sum_vars:
@@ -41,3 +41,9 @@ tongfen_ca_census_ct(
optional parameter to remove NA values when summing, default =
\`TRUE\`
+
+## Value
+
+\`data2\` with the variables in \`data2_sum_vars\` aggregated to a
+common geography matching \`data1\`, identified by the \`GeoUID\` of
+\`data1\`
diff --git a/docs/reference/tongfen_detect_anomalies.html b/docs/reference/tongfen_detect_anomalies.html
new file mode 100644
index 0000000..cc19778
--- /dev/null
+++ b/docs/reference/tongfen_detect_anomalies.html
@@ -0,0 +1,228 @@
+
+Detect likely geocoding anomalies in timelines on a common geography — tongfen_detect_anomalies • tongfen
+ Skip to contents
+
+
+
+
+
+
Detect likely geocoding anomalies in timelines on a common geography
TongFen is only as good as the geocoding that assigned the underlying data to geographic regions
+in the first place. Geocoding varies over time, and the same dwelling units, and the people living
+in them, can get assigned to different neighbouring regions in different years. In a timeline
+on a common geography this shows up as a surprising drop in one region that is offset by
+a corresponding jump in a neighbouring region.
+
This function lists the candidate regions with surprising drops in the given count variable,
+together with the neighbouring region that takes away most of the surprise when both are joined.
+Use it to check for possible problems and to calibrate the parameters before joining regions with
+`tongfen_anomaly_joins`. Not all surprising drops are due to geocoding problems, a drop that is
+not complemented by a neighbouring region is likely real.
+
The surprise of a change between two consecutive years ranges from 0 to 1. Only decreases are surprising,
+the surprise is the product of the surprise of the relative and of the absolute decrease, so that it takes
+a decrease that is large in both relative and absolute terms to be surprising. The total surprise of a
+region is the `p`-norm of the surprises across all changes in the timeline.
+
A candidate region and its neighbour are flagged for joining if joining reduces the total surprise of the
+candidate region to below `cutoff_fact` times its total surprise, or reduces it by more than
+`surprise_reduction_const`, or reduces it to below `cutoff_fact * sum_fact` times the sum of
+the total surprises of both regions. To keep this comparable the surprise after joining is computed
+from the change of the joined regions relative to the counts of the candidate region alone.
data on a common geography, with one row per region, for example as returned by
+`tongfen_aggregate` or `get_tongfen_ca_census`. Needs to be of class sf unless `neighbours` is specified.
+
+
+
variables
+
names of the columns holding the timeline of a count variable like population or
+dwellings, in temporal order. Changes from or to a missing value are not surprising and don't make up for
+surprising changes in neighbouring regions, and joined regions are missing a value if one of the regions they
+are made up of is. Replace missing values by zero beforehand if they stand for regions where nothing got counted
+
+
+
id
+
name of the column that uniquely identifies the regions, default is "TongfenID"
+
+
+
neighbours
+
optional, neighbouring regions as a table with the identifiers of pairs of neighbouring
+regions in the first two columns, or as a neighbours list like the ones returned by `spdep::poly2nb`.
+By default all regions with intersecting geometries are neighbours, which can miss neighbours
+if the geometries have been simplified and don't share their boundaries any more.
+
+
+
rel_scale
+
relative decrease that is half way to full surprise, default is `0.25` for a 25% drop
+
+
+
abs_scale
+
absolute decrease that is half way to full surprise, default is `200`
+
+
+
p
+
exponent of the norm used to combine the surprises across the timeline into the total surprise.
+Large values focus on the most surprising change, 1 adds up the surprises of all changes, default is `4`
+
+
+
surprise_cutoff
+
changes with larger surprise count as surprising, default is `0.15`. Only
+regions with at least one surprising change are candidates
+
+
+
total_surprise_cutoff
+
only regions with larger total surprise are candidates, default is `0.75`
+
+
+
cutoff_fact
+
join regions if the total surprise after joining is lower than this share of the
+total surprise of the candidate region, default is `0.6`
+
+
+
surprise_reduction_const
+
join regions if joining lowers the total surprise by more than this,
+default is `0.15`
+
+
+
sum_fact
+
join regions if the total surprise after joining is lower than `cutoff_fact * sum_fact`
+times the sum of the total surprises of both regions, default is `0.7`
+
+
+
+
Value
+
A tibble with one row for each candidate region, most surprising first, with the identifier of the
+region, the number of surprising changes `surprise_count`, the total surprise `surprise_total`,
+the `period` with the most surprising change, the identifier of the `neighbour` that takes away most of
+the surprise, the total surprise `surprise_total_joined` after joining both and `join` indicating if both
+regions qualify to get joined.
+
+
+
+
Examples
+
# Check 2001 through 2021 dissemination area level population timelines in the
+# City of Vancouver for possible geocoding problems
+if(FALSE){# \dontrun{
+datasets<-c("CA01","CA06","CA11","CA16","CA21")
+meta<-meta_for_additive_variables(datasets,"Population")
+data<-get_tongfen_ca_census(regions=list(CSD="5915022"),meta=meta,level="DA",base_geo="CA21")
+
+anomalies<-tongfen_detect_anomalies(data,paste0("Population_",datasets))
+}# }
+
+
+
+
+
+
+
+
+
+
+
+
+
diff --git a/docs/reference/tongfen_detect_anomalies.md b/docs/reference/tongfen_detect_anomalies.md
new file mode 100644
index 0000000..6ba66b3
--- /dev/null
+++ b/docs/reference/tongfen_detect_anomalies.md
@@ -0,0 +1,152 @@
+# Detect likely geocoding anomalies in timelines on a common geography
+
+**\[experimental\]**
+
+TongFen is only as good as the geocoding that assigned the underlying
+data to geographic regions in the first place. Geocoding varies over
+time, and the same dwelling units, and the people living in them, can
+get assigned to different neighbouring regions in different years. In a
+timeline on a common geography this shows up as a surprising drop in one
+region that is offset by a corresponding jump in a neighbouring region.
+
+This function lists the candidate regions with surprising drops in the
+given count variable, together with the neighbouring region that takes
+away most of the surprise when both are joined. Use it to check for
+possible problems and to calibrate the parameters before joining regions
+with \`tongfen_anomaly_joins\`. Not all surprising drops are due to
+geocoding problems, a drop that is not complemented by a neighbouring
+region is likely real.
+
+The surprise of a change between two consecutive years ranges from 0
+to 1. Only decreases are surprising, the surprise is the product of the
+surprise of the relative and of the absolute decrease, so that it takes
+a decrease that is large in both relative and absolute terms to be
+surprising. The total surprise of a region is the \`p\`-norm of the
+surprises across all changes in the timeline.
+
+A candidate region and its neighbour are flagged for joining if joining
+reduces the total surprise of the candidate region to below
+\`cutoff_fact\` times its total surprise, or reduces it by more than
+\`surprise_reduction_const\`, or reduces it to below \`cutoff_fact \*
+sum_fact\` times the sum of the total surprises of both regions. To keep
+this comparable the surprise after joining is computed from the change
+of the joined regions relative to the counts of the candidate region
+alone.
+
+## Usage
+
+``` r
+tongfen_detect_anomalies(
+ data,
+ variables,
+ id = "TongfenID",
+ neighbours = NULL,
+ rel_scale = 0.25,
+ abs_scale = 200,
+ p = 4,
+ surprise_cutoff = 0.15,
+ total_surprise_cutoff = 0.75,
+ cutoff_fact = 0.6,
+ surprise_reduction_const = 0.15,
+ sum_fact = 0.7
+)
+```
+
+## Arguments
+
+- data:
+
+ data on a common geography, with one row per region, for example as
+ returned by \`tongfen_aggregate\` or \`get_tongfen_ca_census\`. Needs
+ to be of class sf unless \`neighbours\` is specified.
+
+- variables:
+
+ names of the columns holding the timeline of a count variable like
+ population or dwellings, in temporal order. Changes from or to a
+ missing value are not surprising and don't make up for surprising
+ changes in neighbouring regions, and joined regions are missing a
+ value if one of the regions they are made up of is. Replace missing
+ values by zero beforehand if they stand for regions where nothing got
+ counted
+
+- id:
+
+ name of the column that uniquely identifies the regions, default is
+ "TongfenID"
+
+- neighbours:
+
+ optional, neighbouring regions as a table with the identifiers of
+ pairs of neighbouring regions in the first two columns, or as a
+ neighbours list like the ones returned by \`spdep::poly2nb\`. By
+ default all regions with intersecting geometries are neighbours, which
+ can miss neighbours if the geometries have been simplified and don't
+ share their boundaries any more.
+
+- rel_scale:
+
+ relative decrease that is half way to full surprise, default is
+ \`0.25\` for a 25% drop
+
+- abs_scale:
+
+ absolute decrease that is half way to full surprise, default is
+ \`200\`
+
+- p:
+
+ exponent of the norm used to combine the surprises across the timeline
+ into the total surprise. Large values focus on the most surprising
+ change, 1 adds up the surprises of all changes, default is \`4\`
+
+- surprise_cutoff:
+
+ changes with larger surprise count as surprising, default is \`0.15\`.
+ Only regions with at least one surprising change are candidates
+
+- total_surprise_cutoff:
+
+ only regions with larger total surprise are candidates, default is
+ \`0.75\`
+
+- cutoff_fact:
+
+ join regions if the total surprise after joining is lower than this
+ share of the total surprise of the candidate region, default is
+ \`0.6\`
+
+- surprise_reduction_const:
+
+ join regions if joining lowers the total surprise by more than this,
+ default is \`0.15\`
+
+- sum_fact:
+
+ join regions if the total surprise after joining is lower than
+ \`cutoff_fact \* sum_fact\` times the sum of the total surprises of
+ both regions, default is \`0.7\`
+
+## Value
+
+A tibble with one row for each candidate region, most surprising first,
+with the identifier of the region, the number of surprising changes
+\`surprise_count\`, the total surprise \`surprise_total\`, the
+\`period\` with the most surprising change, the identifier of the
+\`neighbour\` that takes away most of the surprise, the total surprise
+\`surprise_total_joined\` after joining both and \`join\` indicating if
+both regions qualify to get joined.
+
+## Examples
+
+``` r
+# Check 2001 through 2021 dissemination area level population timelines in the
+# City of Vancouver for possible geocoding problems
+if (FALSE) { # \dontrun{
+datasets <- c("CA01","CA06","CA11","CA16","CA21")
+meta <- meta_for_additive_variables(datasets,"Population")
+data <- get_tongfen_ca_census(regions=list(CSD="5915022"),meta=meta,level="DA",base_geo="CA21")
+
+anomalies <- tongfen_detect_anomalies(data,paste0("Population_",datasets))
+} # }
+```
diff --git a/docs/reference/tongfen_estimate.html b/docs/reference/tongfen_estimate.html
index e9eeac4..603f34c 100644
--- a/docs/reference/tongfen_estimate.html
+++ b/docs/reference/tongfen_estimate.html
@@ -13,7 +13,7 @@
tongfen
- 0.3.8
+ 0.3.9
@@ -29,6 +29,7 @@
`target` with estimated quantities from `source` as specified by `meta`
+
`target` with estimated quantities from `source` as specified by `meta`, regions in `target`
+that don't overlap with `source` have `NA` values. Columns in `target` can't have the same name
+as the variables to be estimated.
diff --git a/docs/reference/tongfen_estimate.md b/docs/reference/tongfen_estimate.md
index 650af5f..0f17585 100644
--- a/docs/reference/tongfen_estimate.md
+++ b/docs/reference/tongfen_estimate.md
@@ -37,7 +37,9 @@ tongfen_estimate(target, source, meta, na.rm = FALSE)
## Value
\`target\` with estimated quantities from \`source\` as specified by
-\`meta\`
+\`meta\`, regions in \`target\` that don't overlap with \`source\` have
+\`NA\` values. Columns in \`target\` can't have the same name as the
+variables to be estimated.
## Examples
diff --git a/docs/reference/tongfen_estimate_ca_census.html b/docs/reference/tongfen_estimate_ca_census.html
index 2275d36..24c832e 100644
--- a/docs/reference/tongfen_estimate_ca_census.html
+++ b/docs/reference/tongfen_estimate_ca_census.html
@@ -15,7 +15,7 @@
tongfen
- 0.3.8
+ 0.3.9
@@ -31,6 +31,7 @@
`geometry` with the estimated values for the census variables specified by `meta`
+
Examples
-
# Estimate a common geography for 2006 and 2016 dissemination areas in the City of Vancouver
-# based on the geographic data and check estimation errors
+
# Estimate the 2016 population within 1 km of Toronto City Hall from dissemination area level
+# census dataif(FALSE){# \dontrun{toronto_city_hall<-sf::st_point(c(-79.3839,43.6534))%>%sf::st_sfc(crs=4326)%>%
@@ -128,7 +133,7 @@
diff --git a/docs/reference/tongfen_estimate_ca_census.md b/docs/reference/tongfen_estimate_ca_census.md
index fdbddd7..c7bb805 100644
--- a/docs/reference/tongfen_estimate_ca_census.md
+++ b/docs/reference/tongfen_estimate_ca_census.md
@@ -69,11 +69,16 @@ tongfen_estimate_ca_census(
suppress progress messages
+## Value
+
+\`geometry\` with the estimated values for the census variables
+specified by \`meta\`
+
## Examples
``` r
-# Estimate a common geography for 2006 and 2016 dissemination areas in the City of Vancouver
-# based on the geographic data and check estimation errors
+# Estimate the 2016 population within 1 km of Toronto City Hall from dissemination area level
+# census data
if (FALSE) { # \dontrun{
toronto_city_hall <- sf::st_point(c(-79.3839,43.6534)) %>%
sf::st_sfc(crs=4326) %>%
@@ -86,7 +91,7 @@ meta <- meta_for_additive_variables("CA16","Population")
data <- tongfen_estimate_ca_census(toronto_city_hall,meta,level="DA",intersection_level="CT")
print(paste0("Approximately ",scales::comma(data$Population,accuracy=100),
- " people live within a 1 km radius of Toronto City."))
+ " people live within a 1 km radius of Toronto City Hall."))
} # }
```
diff --git a/docs/reference/tongfen_join_correspondence.html b/docs/reference/tongfen_join_correspondence.html
new file mode 100644
index 0000000..10dc817
--- /dev/null
+++ b/docs/reference/tongfen_join_correspondence.html
@@ -0,0 +1,122 @@
+
+Join regions in a correspondence — tongfen_join_correspondence • tongfen
+ Skip to contents
+
+
+
Updates a correspondence so that the given regions are joined, for example to correct for likely geocoding
+anomalies as determined by `tongfen_anomaly_joins`. The updated correspondence can be used in `tongfen_aggregate`
+to aggregate data on the coarser common geography, which works for all variables `tongfen_aggregate` can
+deal with and for data that was not part of detecting the anomalies.
correspondence table with columns the unique geographic identifiers for each of the
+geographies and the TongfenID and TongfenUID, as for example returned by `estimate_tongfen_correspondence`
+or `get_tongfen_correspondence_ca_census`
+
+
+
joins
+
table with the regions to join as returned by `tongfen_anomaly_joins`, with columns
+`TongfenID` and `TongfenID_joined`
+
+
+
+
Value
+
The correspondence with updated TongfenID and TongfenUID for the regions that got joined. If the
+correspondence has a TongfenMethod column "anomaly" gets added to the method of the regions that got joined.
+
+
+
+
Examples
+
# Correct for likely geocoding problems in dissemination area level population timelines
+# and use the updated correspondence to aggregate data on the corrected common geography
+if(FALSE){# \dontrun{
+regions<-list(CSD="5915022")
+datasets<-c("CA01","CA06","CA11","CA16","CA21")
+meta<-meta_for_additive_variables(datasets,"Population")
+data<-get_tongfen_ca_census(regions=regions,meta=meta,level="DA",base_geo="CA21")
+joins<-tongfen_anomaly_joins(data,paste0("Population_",datasets))
+
+correspondence<-get_tongfen_correspondence_ca_census(geo_datasets=datasets,
+ regions=regions,level="DA")%>%
+tongfen_join_correspondence(joins)
+}# }
+
+
+
+
+
+
+
+
+
+
+
+
+
diff --git a/docs/reference/tongfen_join_correspondence.md b/docs/reference/tongfen_join_correspondence.md
new file mode 100644
index 0000000..772d411
--- /dev/null
+++ b/docs/reference/tongfen_join_correspondence.md
@@ -0,0 +1,55 @@
+# Join regions in a correspondence
+
+**\[experimental\]**
+
+Updates a correspondence so that the given regions are joined, for
+example to correct for likely geocoding anomalies as determined by
+\`tongfen_anomaly_joins\`. The updated correspondence can be used in
+\`tongfen_aggregate\` to aggregate data on the coarser common geography,
+which works for all variables \`tongfen_aggregate\` can deal with and
+for data that was not part of detecting the anomalies.
+
+## Usage
+
+``` r
+tongfen_join_correspondence(correspondence, joins)
+```
+
+## Arguments
+
+- correspondence:
+
+ correspondence table with columns the unique geographic identifiers
+ for each of the geographies and the TongfenID and TongfenUID, as for
+ example returned by \`estimate_tongfen_correspondence\` or
+ \`get_tongfen_correspondence_ca_census\`
+
+- joins:
+
+ table with the regions to join as returned by
+ \`tongfen_anomaly_joins\`, with columns \`TongfenID\` and
+ \`TongfenID_joined\`
+
+## Value
+
+The correspondence with updated TongfenID and TongfenUID for the regions
+that got joined. If the correspondence has a TongfenMethod column
+"anomaly" gets added to the method of the regions that got joined.
+
+## Examples
+
+``` r
+# Correct for likely geocoding problems in dissemination area level population timelines
+# and use the updated correspondence to aggregate data on the corrected common geography
+if (FALSE) { # \dontrun{
+regions <- list(CSD="5915022")
+datasets <- c("CA01","CA06","CA11","CA16","CA21")
+meta <- meta_for_additive_variables(datasets,"Population")
+data <- get_tongfen_ca_census(regions=regions,meta=meta,level="DA",base_geo="CA21")
+joins <- tongfen_anomaly_joins(data,paste0("Population_",datasets))
+
+correspondence <- get_tongfen_correspondence_ca_census(geo_datasets=datasets,
+ regions=regions,level="DA") %>%
+ tongfen_join_correspondence(joins)
+} # }
+```
diff --git a/docs/reference/tongfen_join_regions.html b/docs/reference/tongfen_join_regions.html
new file mode 100644
index 0000000..4aff240
--- /dev/null
+++ b/docs/reference/tongfen_join_regions.html
@@ -0,0 +1,144 @@
+
+Join regions in data on a common geography — tongfen_join_regions • tongfen
+ Skip to contents
+
+
+
Joins regions in data that has already been aggregated to a common geography, for example to correct
+for likely geocoding anomalies as determined by `tongfen_anomaly_joins`. The data, and the geometries if
+the data is of class sf, of the regions that get joined are aggregated, all other regions are left as they are.
+
Variables are aggregated according to the metadata, numeric variables that are not part of the metadata
+are assumed to be additive. Variables that are not additive, like averages, can only be aggregated if their
+parent variable is part of the data. If that is not the case use `tongfen_join_correspondence` to update the
+correspondence the data was built from and aggregate the original data again with `tongfen_aggregate`.
+
+
+
+
Usage
+
tongfen_join_regions(data, joins, meta =NULL, id ="TongfenID", na.rm =TRUE)
+
+
+
+
Arguments
+
+
+
data
+
data on a common geography, with one row per region, for example as returned by
+`tongfen_aggregate` or `get_tongfen_ca_census`
+
+
+
joins
+
table with the regions to join as returned by `tongfen_anomaly_joins`, with the identifier
+of the region and the identifier of the joined region it becomes part of in the column named like the
+identifier with suffix `_joined`
+
+
+
meta
+
optional metadata containing aggregation rules as for example returned by `meta_for_ca_census_vectors`,
+variables are matched by their label. Numeric variables that are not part of the metadata are treated as
+additive, if `NULL` (the default) that is the case for all numeric variables
+
+
+
id
+
name of the column that uniquely identifies the regions, default is "TongfenID"
+
+
+
na.rm
+
logical, determines how NA values should be treated when aggregating variables,
+default is `TRUE`
+
+
+
+
Value
+
The data with the regions joined. Joined regions take the place and the identifier of the
+region with the smallest identifier among the regions they are made up of. Variables that are
+not numeric and not part of the metadata are `NA` for joined regions.
+
+
+
+
Examples
+
# Correct 2001 through 2021 dissemination area level population timelines in the
+# City of Vancouver for likely geocoding problems
+if(FALSE){# \dontrun{
+datasets<-c("CA01","CA06","CA11","CA16","CA21")
+meta<-meta_for_additive_variables(datasets,"Population")
+data<-get_tongfen_ca_census(regions=list(CSD="5915022"),meta=meta,level="DA",base_geo="CA21")
+
+joins<-tongfen_anomaly_joins(data,paste0("Population_",datasets))
+corrected_data<-tongfen_join_regions(data,joins,meta)
+}# }
+
+
+
+
+
+
+
+
+
+
+
+
+
diff --git a/docs/reference/tongfen_join_regions.md b/docs/reference/tongfen_join_regions.md
new file mode 100644
index 0000000..40d63d0
--- /dev/null
+++ b/docs/reference/tongfen_join_regions.md
@@ -0,0 +1,77 @@
+# Join regions in data on a common geography
+
+**\[experimental\]**
+
+Joins regions in data that has already been aggregated to a common
+geography, for example to correct for likely geocoding anomalies as
+determined by \`tongfen_anomaly_joins\`. The data, and the geometries if
+the data is of class sf, of the regions that get joined are aggregated,
+all other regions are left as they are.
+
+Variables are aggregated according to the metadata, numeric variables
+that are not part of the metadata are assumed to be additive. Variables
+that are not additive, like averages, can only be aggregated if their
+parent variable is part of the data. If that is not the case use
+\`tongfen_join_correspondence\` to update the correspondence the data
+was built from and aggregate the original data again with
+\`tongfen_aggregate\`.
+
+## Usage
+
+``` r
+tongfen_join_regions(data, joins, meta = NULL, id = "TongfenID", na.rm = TRUE)
+```
+
+## Arguments
+
+- data:
+
+ data on a common geography, with one row per region, for example as
+ returned by \`tongfen_aggregate\` or \`get_tongfen_ca_census\`
+
+- joins:
+
+ table with the regions to join as returned by
+ \`tongfen_anomaly_joins\`, with the identifier of the region and the
+ identifier of the joined region it becomes part of in the column named
+ like the identifier with suffix \`\_joined\`
+
+- meta:
+
+ optional metadata containing aggregation rules as for example returned
+ by \`meta_for_ca_census_vectors\`, variables are matched by their
+ label. Numeric variables that are not part of the metadata are treated
+ as additive, if \`NULL\` (the default) that is the case for all
+ numeric variables
+
+- id:
+
+ name of the column that uniquely identifies the regions, default is
+ "TongfenID"
+
+- na.rm:
+
+ logical, determines how NA values should be treated when aggregating
+ variables, default is \`TRUE\`
+
+## Value
+
+The data with the regions joined. Joined regions take the place and the
+identifier of the region with the smallest identifier among the regions
+they are made up of. Variables that are not numeric and not part of the
+metadata are \`NA\` for joined regions.
+
+## Examples
+
+``` r
+# Correct 2001 through 2021 dissemination area level population timelines in the
+# City of Vancouver for likely geocoding problems
+if (FALSE) { # \dontrun{
+datasets <- c("CA01","CA06","CA11","CA16","CA21")
+meta <- meta_for_additive_variables(datasets,"Population")
+data <- get_tongfen_ca_census(regions=list(CSD="5915022"),meta=meta,level="DA",base_geo="CA21")
+
+joins <- tongfen_anomaly_joins(data,paste0("Population_",datasets))
+corrected_data <- tongfen_join_regions(data,joins,meta)
+} # }
+```
diff --git a/docs/reference/tongfen_tag_largest_overlap.html b/docs/reference/tongfen_tag_largest_overlap.html
index c580c9c..4d4ff1a 100644
--- a/docs/reference/tongfen_tag_largest_overlap.html
+++ b/docs/reference/tongfen_tag_largest_overlap.html
@@ -9,7 +9,7 @@
tongfen
- 0.3.8
+ 0.3.9
@@ -25,6 +25,7 @@
`source` with extra column with name `"target_id"` and column `...overlap_fraction` with
-the proportion of overlap of the target geometry with the respective `target_id`
+
`source` with extra column with the name given by `target_id` and column `...overlap_fraction` with
+the proportion of the area of the source region that overlaps with the region in `target` with that id
Examples
-
# Estimate 2006 Populatino in the City of Vancouver dissemination ares on 2016 census geoographies
+
# Tag 2016 dissemination areas in the City of Vancouver by the 2006 census tract they overlap
+# the most withif(FALSE){# \dontrun{
-geo1<-cancensus::get_census("CA06",regions=list(CSD="5915022"),geo_format='sf',level='DA')
+geo1<-cancensus::get_census("CA06",regions=list(CSD="5915022"),geo_format='sf',level='CT')geo2<-cancensus::get_census("CA16",regions=list(CSD="5915022"),geo_format='sf',level='DA')
-meta<-meta_for_additive_variables("CA06","Population")
-result<-tongfen_estimate(geo2%>%rename(Population_2016=Population),geo1,meta)
+result<-tongfen_tag_largest_overlap(geo2,geo1%>%select(CT_2006=GeoUID),"CT_2006")}# }
diff --git a/docs/reference/tongfen_tag_largest_overlap.md b/docs/reference/tongfen_tag_largest_overlap.md
index c400728..85f25f2 100644
--- a/docs/reference/tongfen_tag_largest_overlap.md
+++ b/docs/reference/tongfen_tag_largest_overlap.md
@@ -27,18 +27,18 @@ tongfen_tag_largest_overlap(source, target, target_id)
## Value
-\`source\` with extra column with name \`"target_id"\` and column
-\`...overlap_fraction\` with the proportion of overlap of the target
-geometry with the respective \`target_id\`
+\`source\` with extra column with the name given by \`target_id\` and
+column \`...overlap_fraction\` with the proportion of the area of the
+source region that overlaps with the region in \`target\` with that id
## Examples
``` r
-# Estimate 2006 Populatino in the City of Vancouver dissemination ares on 2016 census geoographies
+# Tag 2016 dissemination areas in the City of Vancouver by the 2006 census tract they overlap
+# the most with
if (FALSE) { # \dontrun{
-geo1 <- cancensus::get_census("CA06",regions=list(CSD="5915022"),geo_format='sf',level='DA')
+geo1 <- cancensus::get_census("CA06",regions=list(CSD="5915022"),geo_format='sf',level='CT')
geo2 <- cancensus::get_census("CA16",regions=list(CSD="5915022"),geo_format='sf',level='DA')
-meta <- meta_for_additive_variables("CA06","Population")
-result <- tongfen_estimate(geo2 %>% rename(Population_2016=Population),geo1,meta)
+result <- tongfen_tag_largest_overlap(geo2,geo1 %>% select(CT_2006=GeoUID),"CT_2006")
} # }
```
diff --git a/docs/reference/vancouver_elections_data_2015.html b/docs/reference/vancouver_elections_data_2015.html
index 1e4db1c..9807c9b 100644
--- a/docs/reference/vancouver_elections_data_2015.html
+++ b/docs/reference/vancouver_elections_data_2015.html
@@ -7,7 +7,7 @@
tongfen
- 0.3.8
+ 0.3.9
@@ -23,6 +23,7 @@
Author<
diff --git a/docs/search.json b/docs/search.json
index ca37f24..b1f9d5e 100644
--- a/docs/search.json
+++ b/docs/search.json
@@ -1 +1 @@
-[{"path":"https://mountainmath.github.io/tongfen/LICENSE.html","id":null,"dir":"","previous_headings":"","what":"MIT License","title":"MIT License","text":"Copyright (c) 2020 Jens von Bergmann Permission hereby granted, free charge, person obtaining copy software associated documentation files (“Software”), deal Software without restriction, including without limitation rights use, copy, modify, merge, publish, distribute, sublicense, /sell copies Software, permit persons Software furnished , subject following conditions: copyright notice permission notice shall included copies substantial portions Software. SOFTWARE PROVIDED “”, WITHOUT WARRANTY KIND, EXPRESS IMPLIED, INCLUDING LIMITED WARRANTIES MERCHANTABILITY, FITNESS PARTICULAR PURPOSE NONINFRINGEMENT. EVENT SHALL AUTHORS COPYRIGHT HOLDERS LIABLE CLAIM, DAMAGES LIABILITY, WHETHER ACTION CONTRACT, TORT OTHERWISE, ARISING , CONNECTION SOFTWARE USE DEALINGS SOFTWARE.","code":""},{"path":"https://mountainmath.github.io/tongfen/articles/tongfen.html","id":"population-change","dir":"Articles","previous_headings":"","what":"Population change","title":"General TongFen","text":"’s time go back original goal mapping population change. need specify aggregate population data, simply adding . meta_for_additive_variables convenience function generates appropriate metatdata specifies deal data. ’s left add population data. choose 2001 base year clipped boundaries look better.","code":"meta <- meta_for_additive_variables(years,\"Population\") meta #> # A tibble: 4 × 8 #> variable dataset label type aggregation rule geo_dataset parent #> #> 1 Population 2001 Population_2001 Manual Additive Addi… 2001 NA #> 2 Population 2006 Population_2006 Manual Additive Addi… 2006 NA #> 3 Population 2011 Population_2011 Manual Additive Addi… 2011 NA #> 4 Population 2016 Population_2016 Manual Additive Addi… 2016 NA breaks = c(-0.15,-0.1,-0.075,-0.05,-0.025,0,0.025,0.05,0.1,0.2,0.3) labels = c(\"-15% to -10%\",\"-10% to -7.5%\",\"-7.5% to -5%\",\"-5% to -2.5%\",\"-2.5% to 0%\",\"0% to 2.5%\",\"2.5% to 5%\",\"5% to 10%\",\"10% to 20%\",\"20% to 30%\") colors <- RColorBrewer::brewer.pal(10,\"PiYG\") compute_population_change_metrics <- function(data) { geometric_average <- function(x,n){sign(x) * (exp(log(1+abs(x))/n)-1)} data %>% mutate(`2001 - 2006`=geometric_average((`Population_2006`-`Population_2001`)/`Population_2001`,5), `2006 - 2011`=geometric_average((`Population_2011`-`Population_2006`)/`Population_2006`,5), `2011 - 2016`=geometric_average((`Population_2016`-`Population_2011`)/`Population_2011`,5), `2001 - 2016`=geometric_average((`Population_2016`-`Population_2001`)/`Population_2001`,15)) %>% gather(key=\"Period\",value=\"Population Change\",c(\"2001 - 2006\",\"2006 - 2011\",\"2011 - 2016\",\"2001 - 2016\")) %>% mutate(Period=factor(Period,levels=c(\"2001 - 2006\",\"2006 - 2011\",\"2011 - 2016\",\"2001 - 2016\"))) %>% mutate(c=cut(`Population Change`,breaks=breaks, labels=labels)) } plot_data <- tongfen_aggregate(data,correspondence,meta=meta,base_geo = \"2001\") %>% compute_population_change_metrics() ggplot(plot_data,aes(fill=c)) + geom_sf(size=0.1) + scale_fill_manual(values=setNames(colors,labels)) + facet_wrap(\"Period\",ncol=2) + coord_sf(datum=NA) + labs(fill=\"Average Annual\\nPopulation Change\", title=\"Vancouver population change\", caption = \"StatCan Census 2001-2016\")"},{"path":"https://mountainmath.github.io/tongfen/articles/tongfen_ca.html","id":"aggregating-up-data-across-regions","dir":"Articles","previous_headings":"","what":"Aggregating up data across regions","title":"TongFen for Canadian census data","text":"Another application simply aggregating variables selection regions single census. Suppose want understand share renters Vancouver School District, well share renter households spending 30% income housing. 53% Vancouver School District households rent, 45% shelter cost burdened.","code":"vectors <- c(\"v_CA16_4836\",\"v_CA16_4838\",\"v_CA16_4899\") meta=meta_for_ca_census_vectors(vectors) %>% bind_rows(meta_for_additive_variables(\"CA16\",c(\"Population\",\"Dwellings\",\"Households\"))) vsb_regions <- list(CSD=c(\"5915022\",\"5915803\"), CT=c(\"9330069.01\",\"9330069.02\",\"9330069.00\")) vsb <- get_census(\"CA16\",regions=vsb_regions,vectors=meta$variable,labels=\"short\") vsb <- aggregate_data_with_meta(vsb, meta) %>% mutate(Total=v_CA16_4836,Renters=v_CA16_4838,rent_poor=v_CA16_4899/100) %>% mutate(rent_share=Renters/Total)"},{"path":"https://mountainmath.github.io/tongfen/articles/tongfen_ca.html","id":"change-in-vancouver-children","dir":"Articles","previous_headings":"Aggregating up data across regions","what":"Change in Vancouver children","title":"TongFen for Canadian census data","text":"Data can also obtained dissemination area level. example look change children aged 0 14 Vancouver School District.","code":"variables <- c(\"2016_0-14\"=\"v_CA16_4\", \"2011_0-4\"=\"v_CA11F_8\",\"2011_5-9\"=\"v_CA11F_11\",\"2011_10-14\"=\"v_CA11F_14\") meta <- meta_for_ca_census_vectors(variables) %>% bind_rows(meta_for_additive_variables(c(\"CA11\",\"CA16\"),\"Population\")) children_data <- get_tongfen_ca_census(regions = vsb_regions, meta = meta, level=\"DA\", base_geo = \"CA16\", quiet = TRUE) %>% mutate(`2011_0-14`=purrr::reduce(select(sf::st_set_geometry(.,NULL), starts_with(\"2011_\")), `+`)) %>% mutate(change=`2016_0-14`/Population_CA16-`2011_0-14`/Population_CA11) ggplot(children_data,aes(fill=change)) + geom_sf(size=0.1) + scale_fill_gradient2(labels=scales::percent) + coord_sf(datum=NA) + labs(title=\"Percentage point change in share of children aged 0-14 between 2011 and 2016\",fill=NULL)"},{"path":"https://mountainmath.github.io/tongfen/articles/tongfen_us.html","id":"bridging-to-the-2020-census","dir":"Articles","previous_headings":"","what":"Bridging to the 2020 census","title":"TongFen for US census data","text":"works across 2010 2020 censuses. Two things change. 2020 census renamed variables, population occupied housing units H8_001N households H3_002N. live Demographic Housing Characteristics file, whereas tidycensus reads PL 94-171 redistricting file 2020 default, point right one via sumfile. takes single value censuses, one named dataset .","code":"meta_2020 <- bind_rows( meta_for_additive_variables(\"dec2010\",c(population_2010=\"H011001\", households_2010=\"H013001\")), meta_for_additive_variables(\"dec2020\",c(population_2020=\"H8_001N\", households_2020=\"H3_002N\"))) census_data_2020 <- get_tongfen_us_census(regions = list(state=\"CA\"), meta=meta_2020, level=\"tract\", sumfile=c(dec2020=\"dhc\")) %>% mutate(change=population_2020/households_2020-population_2010/households_2010) census_data_2020 %>% mutate(c=cut(change,c(-Inf,-0.5,-0.3,-0.2,-0.1,0,0.1,0.2,0.3,0.5,Inf))) %>% ggplot() + geom_sf(aes(fill=c), size=0.05) + scale_fill_brewer(palette = \"RdYlGn\") + labs(title=\"Bay area change in average household size 2010-2020\", fill=NULL) + coord_sf(datum=NA,xlim=c(-122.6,-121.7),ylim=c(37.2,37.9))"},{"path":"https://mountainmath.github.io/tongfen/articles/tongfen_us.html","id":"notes-on-the-common-geography","dir":"Articles","previous_headings":"","what":"Notes on the common geography","title":"TongFen for US census data","text":"Census Bureau relationship files correspondences built geometric overlays, list every place two censuses’ geographies intersect, including slivers along boundaries shifted metres. Chaining together merges regions nothing , min_area_share sets much area two regions common count related. default 0.01 works well, raising gives finer common geographies risk separating regions genuinely change. region ever dropped, region’s parts fall cutoff largest part kept. Correspondence tables reach back one census data . Census Bureau retired 1990 API endpoint, tidycensus fetch 1990 data, get_tongfen_correspondence_us_census match 1990 tracts later censuses. use , get 1990 data elsewhere, example NHGIS via ipumsr package, hand tongfen_aggregate along correspondence table.","code":""},{"path":"https://mountainmath.github.io/tongfen/authors.html","id":null,"dir":"","previous_headings":"","what":"Authors","title":"Authors and Citation","text":"Jens von Bergmann. Author, maintainer. creator maintainer","code":""},{"path":"https://mountainmath.github.io/tongfen/authors.html","id":"citation","dir":"","previous_headings":"","what":"Citation","title":"Authors and Citation","text":"von Bergmann J (2026). tongfen: Make Data Based Different Geographies Comparable. R package version 0.3.8, https://github.com/mountainMath/tongfen.","code":"@Manual{, title = {tongfen: Make Data Based on Different Geographies Comparable}, author = {Jens {von Bergmann}}, year = {2026}, note = {R package version 0.3.8}, url = {https://github.com/mountainMath/tongfen}, }"},{"path":"https://mountainmath.github.io/tongfen/index.html","id":"tongfen","dir":"","previous_headings":"","what":"Make Data Based on Different Geographies Comparable","title":"Make Data Based on Different Geographies Comparable","text":"TongFen (通分) means convert two fractions least common denominator, typically preparation manipulation like addition subtraction. English, ’s mouthful sounds complicated. Chinese word , TongFen, makes process appear simple. working geospatial datasets often want compare data given different regions. example census data election data. data two different censuses. properly compare data first need convert common geography. process quite analogous process TongFen fractions, appropriate term give simple name. Using tongfen package, preparing data disparate geographies comparison converting common geography easy typing tongfen.","code":""},{"path":"https://mountainmath.github.io/tongfen/index.html","id":"reference","dir":"","previous_headings":"","what":"Reference","title":"Make Data Based on Different Geographies Comparable","text":"TongFen home page reference guide","code":""},{"path":"https://mountainmath.github.io/tongfen/index.html","id":"installing-the-package","dir":"","previous_headings":"","what":"Installing the package","title":"Make Data Based on Different Geographies Comparable","text":"latest development version can installed GitHub.","code":"install.packages(\"tongfen\") remotes::install_github(\"mountainmath/tongfen\") library(tongfen)"},{"path":"https://mountainmath.github.io/tongfen/index.html","id":"caching-correspondence-files","dir":"","previous_headings":"","what":"Caching correspondence files","title":"Make Data Based on Different Geographies Comparable","text":"get_tongfen_census_ct get_tongfen_census_ct_from_da methods make use StatCan correspondence files. speed process useful permanently cache files instead download repeatedly. caching desired, set either options(\"tongfen.cache_path\"=\"\") Sys.setenv(\"tongfen.cache_path\"=\"\") options(\"custom_data_path\"=\"\") .Rprofile .Renviron file.","code":""},{"path":"https://mountainmath.github.io/tongfen/index.html","id":"general-tongfen","dir":"","previous_headings":"","what":"General TongFen","title":"Make Data Based on Different Geographies Comparable","text":"tongfen package build around following basic TongFen workflow: Given list datasets diverse geographies, generate correspondence table links geographies specifies aggregate (least) common geography via estimate_tongfen_correspondence. generate metadata specifies variables can aggregated , meta_for_additive_variables function additive variables. Use correspondence table metadata generate dataset variables original datasets aggregated common geography via tongfen_aggregate. convenience function validate geographic TongFen fit via area comparison available via check_tongfen_areas, allows explore deal spatial mismatches TongFen.","code":""},{"path":"https://mountainmath.github.io/tongfen/index.html","id":"aggregation-of-variables","dir":"","previous_headings":"General TongFen","what":"Aggregation of variables","title":"Make Data Based on Different Geographies Comparable","text":"Finding common tiling several different yet congruent geographies one part problem TongFen addresses, aggregating variables part. tongfen package deals using metadata table specifies variables aggregated. ’s simplest form values simply added . meta_for_additive_variables convenience function builds metadata additive variables. Metadata non-additive variables like averages, ratios percentages needs care build, requires additional information parent variable specifies denominator average, ratio percentage. data, like medians, can’t aggregated , although tongfen can provide estimates medians aggregated geographies treating averages.","code":""},{"path":"https://mountainmath.github.io/tongfen/index.html","id":"packaged-data","dir":"","previous_headings":"General TongFen","what":"Packaged data","title":"Make Data Based on Different Geographies Comparable","text":"package ships subset voting data Elections Canada 42nd 43rd federal elections well polling district geographies 42nd 43rd. facilitates running example vignette polling districts without download external data. available open data covered Open Government Licence - Canda.","code":""},{"path":"https://mountainmath.github.io/tongfen/index.html","id":"data-specific-implementations","dir":"","previous_headings":"","what":"Data-specific implementations","title":"Make Data Based on Different Geographies Comparable","text":"need TongFen comes frequently certain types geographies. Census geographies one example. cases data sources come correspondence files go beyond geographic matchup also join regions alleviate data integrity problems like geocoding issues. cases can worthwhile wrap data acquisition TongFen one convenience function, also extend TongFen method parameter allow external correspondence files used.","code":""},{"path":"https://mountainmath.github.io/tongfen/index.html","id":"canadian-census-data","dir":"","previous_headings":"Data-specific implementations","what":"Canadian census data","title":"Make Data Based on Different Geographies Comparable","text":"package well-integrated work Canadian census data two essential ways. * meta_for_ca_census_vectors builds rich metadata given list Canadian census variables utilizing metadata available via CensusMapper. particular, automates proper aggregation non-count variables like averages, ratios percentages. * get_tongfen_ca_census wraps process data acquisition (via CensusMapper cancensus package tongfen one convenience function. time adds TongFen method = \"statcan\" option uses Statistics Canada correspondence files build common geography. * get_tongfen_correspondence_ca_census function breaks correspondence generation aid process accessing Statistics Canada correspondence files (better integration generating correspondences Canadian census geographies general) facilitate mixing non-census data coming census geographies, like example CMHC data.","code":""},{"path":"https://mountainmath.github.io/tongfen/index.html","id":"us-census-data","dir":"","previous_headings":"Data-specific implementations","what":"US census data","title":"Make Data Based on Different Geographies Comparable","text":"get_tongfen_us_census integrates data acquisition (via tidycensus package) TongFen, adds tongfen method = \"census.gov\" use US Census Bureau correspondence files matching.","code":""},{"path":"https://mountainmath.github.io/tongfen/index.html","id":"other-implementations","dir":"","previous_headings":"","what":"Other implementations","title":"Make Data Based on Different Geographies Comparable","text":"tongfen package open add extensions specialized data sources, well extensions existing ones.","code":""},{"path":"https://mountainmath.github.io/tongfen/index.html","id":"fixed-target-geography-estimation","dir":"","previous_headings":"","what":"Fixed target geography estimation","title":"Make Data Based on Different Geographies Comparable","text":"geographies aren’t sufficiently congruent target geography fixed, won’t able use tongfen methods compute data common geography instead rely estimates. tongfen_estimate makes assumption underlying geographies returns estimates data target geography. uses area-weighted interpolation achieve , can refined dasymmetric estimates using proportional_reaggregate function. method example works independent nature underlying geographies, comes heavy price estimate. useful research purposes also need methods estimate errors introduces effects subsequent analysis results. Methods facilitate still active development.","code":""},{"path":"https://mountainmath.github.io/tongfen/index.html","id":"cite-tongfen","dir":"","previous_headings":"Fixed target geography estimation","what":"Cite tongfen","title":"Make Data Based on Different Geographies Comparable","text":"wish cite tongfen: von Bergmann, J. (2024). tongfen: R package Make Data Based Different Geographies Comparable. v0.3.7. DOI: 10.32614/CRAN.package.tongfen BibTeX entry LaTeX users ","code":"@Manual{tongfen, author = {Jens {von Bergmann}}, title = {tongfen: R package to Make Data Based on Different Geographies Comparable}, year = {2024}, doi = {10.32614/CRAN.package.tongfen}, note = {R package version 0.3.7}, url = {https://mountainmath.github.io/tongfen/}, }"},{"path":"https://mountainmath.github.io/tongfen/reference/add_census_ca_base_variables.html","id":null,"dir":"Reference","previous_headings":"","what":"Generate metadata from Candian census vectors — add_census_ca_base_variables","title":"Generate metadata from Candian census vectors — add_census_ca_base_variables","text":"Add Population, Dwellings, Household counts metadata","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/add_census_ca_base_variables.html","id":"ref-usage","dir":"Reference","previous_headings":"","what":"Usage","title":"Generate metadata from Candian census vectors — add_census_ca_base_variables","text":"","code":"add_census_ca_base_variables(meta)"},{"path":"https://mountainmath.github.io/tongfen/reference/add_census_ca_base_variables.html","id":"arguments","dir":"Reference","previous_headings":"","what":"Arguments","title":"Generate metadata from Candian census vectors — add_census_ca_base_variables","text":"meta ribble metadata example provided `meta_for_ca_census_vectors`","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/add_census_ca_base_variables.html","id":"value","dir":"Reference","previous_headings":"","what":"Value","title":"Generate metadata from Candian census vectors — add_census_ca_base_variables","text":"tibble metadata","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/aggregate_data_with_meta.html","id":null,"dir":"Reference","previous_headings":"","what":"Aggregate variables in grouped data — aggregate_data_with_meta","title":"Aggregate variables in grouped data — aggregate_data_with_meta","text":"Aggregate census data , assumes data grouped aggregation Uses data meta determine aggregate ","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/aggregate_data_with_meta.html","id":"ref-usage","dir":"Reference","previous_headings":"","what":"Usage","title":"Aggregate variables in grouped data — aggregate_data_with_meta","text":"","code":"aggregate_data_with_meta(data, meta, geo = FALSE, na.rm = TRUE, quiet = FALSE)"},{"path":"https://mountainmath.github.io/tongfen/reference/aggregate_data_with_meta.html","id":"arguments","dir":"Reference","previous_headings":"","what":"Arguments","title":"Aggregate variables in grouped data — aggregate_data_with_meta","text":"data census data obtained get_census call, grouped TongfenID meta list variables aggregation information obtained meta_for_vectors geo logical, also aggregate geographic data na.rm logical, NA values ignored carried . quiet logical, emit messages set `TRUE`","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/aggregate_data_with_meta.html","id":"value","dir":"Reference","previous_headings":"","what":"Value","title":"Aggregate variables in grouped data — aggregate_data_with_meta","text":"data frame variables aggregated new common geography","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/aggregate_data_with_meta.html","id":"ref-examples","dir":"Reference","previous_headings":"","what":"Examples","title":"Aggregate variables in grouped data — aggregate_data_with_meta","text":"","code":"# Aggregate population from DA level to grouped by CT_UID if (FALSE) { # \\dontrun{ geo <- cancensus::get_census(\"CA06\",regions=list(CSD=\"5915022\"),level='DA') meta <- meta_for_additive_variables(\"CA06\",\"Population\") result <- aggregate_data_with_meta(geo %>% group_by(CT_UID),meta) } # }"},{"path":"https://mountainmath.github.io/tongfen/reference/check_tongfen_areas.html","id":null,"dir":"Reference","previous_headings":"","what":"Check geographic integrety — check_tongfen_areas","title":"Check geographic integrety — check_tongfen_areas","text":"Sanity check areas estimated tongfen correspondence. useful example total extent geo1 geo2 differ regions edges large difference overlap. result diagnostic, pass/fail test. Geographies different years simplified independently differ water features cut , sizable area mismatch mean regions matched incorrectly.","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/check_tongfen_areas.html","id":"ref-usage","dir":"Reference","previous_headings":"","what":"Usage","title":"Check geographic integrety — check_tongfen_areas","text":"","code":"check_tongfen_areas(data, correspondence)"},{"path":"https://mountainmath.github.io/tongfen/reference/check_tongfen_areas.html","id":"arguments","dir":"Reference","previous_headings":"","what":"Arguments","title":"Check geographic integrety — check_tongfen_areas","text":"data alist geogrpahic data class sf correspondence Correspondence table columns unique geographic identifiers geographies TongfenID (optionally TongfenUID TongfenMethod) returned `estimate_tongfen_correspondence`.","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/check_tongfen_areas.html","id":"value","dir":"Reference","previous_headings":"","what":"Value","title":"Check geographic integrety — check_tongfen_areas","text":"table columns `TongfenID`, geo_identifiers, areas aggregated regions corresponding geographic identifier column, tongfen estimation method maximum log ratio areas.","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/check_tongfen_areas.html","id":"ref-examples","dir":"Reference","previous_headings":"","what":"Examples","title":"Check geographic integrety — check_tongfen_areas","text":"","code":"# Estimate a common geography for 2006 and 2016 dissemination areas in the City of Vancouver # based on the geographic data and check estimation errors if (FALSE) { # \\dontrun{ regions <- list(CSD=\"5915022\") data_06 <- cancensus::get_census(\"CA06\",regions=regions,geo_format='sf',level=\"DA\") %>% rename(GeoUID_06=GeoUID) data_16 <- cancensus::get_census(\"CA16\",regions=regions,geo_format=\"sf\",level=\"DA\") %>% rename(GeoUID_16=GeoUID) correspondence <- estimate_tongfen_correspondence(list(data_06, data_16), c(\"GeoUID_06\",\"GeoUID_16\")) area_check <- check_tongfen_areas(list(data_06, data_16),correspondence) } # }"},{"path":"https://mountainmath.github.io/tongfen/reference/check_tongfen_single_areas.html","id":null,"dir":"Reference","previous_headings":"","what":"Check geographic integrety — check_tongfen_single_areas","title":"Check geographic integrety — check_tongfen_single_areas","text":"Sanity check areas estimated tongfen correspondence. useful example total extent geo1 geo2 differ regions edges large difference overlap.","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/check_tongfen_single_areas.html","id":"ref-usage","dir":"Reference","previous_headings":"","what":"Usage","title":"Check geographic integrety — check_tongfen_single_areas","text":"","code":"check_tongfen_single_areas(geo1, geo2, correspondence)"},{"path":"https://mountainmath.github.io/tongfen/reference/check_tongfen_single_areas.html","id":"arguments","dir":"Reference","previous_headings":"","what":"Arguments","title":"Check geographic integrety — check_tongfen_single_areas","text":"geo1 input geometry 1 class sf geo2 input geometry 2 class sf correspondence Correspondence table `geo1` `geo2` e.g. returned `estimate_tongfen_correspondence`.","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/check_tongfen_single_areas.html","id":"value","dir":"Reference","previous_headings":"","what":"Value","title":"Check geographic integrety — check_tongfen_single_areas","text":"table columns `TongfenID`, `area1` `area2`, row corresponds unique `TongfenID` `correspondence` table columns hold areas regions aggregated `geo1` `geo2`.`","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/estimate_tongfen_correspondence.html","id":null,"dir":"Reference","previous_headings":"","what":"Generate togfen correspondence for list of geographies — estimate_tongfen_correspondence","title":"Generate togfen correspondence for list of geographies — estimate_tongfen_correspondence","text":"Get correspondence data arbitrary congruent geometries. Congruent means one can obtain common tiling aggregating several sub-geometries two input geo data. Worst case scenario common tiling given unioning sub-geometries finer common tiling.","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/estimate_tongfen_correspondence.html","id":"ref-usage","dir":"Reference","previous_headings":"","what":"Usage","title":"Generate togfen correspondence for list of geographies — estimate_tongfen_correspondence","text":"","code":"estimate_tongfen_correspondence( data, geo_identifiers, method = \"estimate\", tolerance = 50, computation_crs = NULL )"},{"path":"https://mountainmath.github.io/tongfen/reference/estimate_tongfen_correspondence.html","id":"arguments","dir":"Reference","previous_headings":"","what":"Arguments","title":"Generate togfen correspondence for list of geographies — estimate_tongfen_correspondence","text":"data list geometries class sf geo_identifiers vector unique geographic identifiers list entry data. method aggregation method. Possible values \"estimate\" \"identifier\". \"estimate\" estimates correspondence purely geographic data. \"identifier\" assumes regions identical geo_identifiers , uses \"estimate\" method remaining regions. Default \"estimate\". tolerance tolerance (projected coordinate units `computation_crs`) feature matching computation_crs optional crs computation carried , defaults crs first entry data parameter.","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/estimate_tongfen_correspondence.html","id":"value","dir":"Reference","previous_headings":"","what":"Value","title":"Generate togfen correspondence for list of geographies — estimate_tongfen_correspondence","text":"correspondence table linking geo1_uid geo2_uid unique TongfenID TongfenUID columns enumerate common geometry.","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/estimate_tongfen_correspondence.html","id":"ref-examples","dir":"Reference","previous_headings":"","what":"Examples","title":"Generate togfen correspondence for list of geographies — estimate_tongfen_correspondence","text":"","code":"# Estimate a common geography for 2006 and 2016 dissemination areas in the City of Vancouver # based on the geographic data. if (FALSE) { # \\dontrun{ regions <- list(CSD=\"5915022\") data_06 <- cancensus::get_census(\"CA06\",regions=regions,geo_format='sf',level=\"DA\") %>% rename(GeoUID_06=GeoUID) data_16 <- cancensus::get_census(\"CA16\",regions=regions,geo_format=\"sf\",level=\"DA\") %>% rename(GeoUID_16=GeoUID) correspondence <- estimate_tongfen_correspondence(list(data_06, data_16), c(\"GeoUID_06\",\"GeoUID_16\")) } # }"},{"path":"https://mountainmath.github.io/tongfen/reference/estimate_tongfen_single_correspondence.html","id":null,"dir":"Reference","previous_headings":"","what":"Generate togfen correspondence for two geographies — estimate_tongfen_single_correspondence","title":"Generate togfen correspondence for two geographies — estimate_tongfen_single_correspondence","text":"Get correspondence data arbitrary congruent geometries. Congruent means one can obtain common tiling aggregating several sub-geometries two input geo data. Worst case scenario common tiling given unioning sub-geometries finer common tiling.","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/estimate_tongfen_single_correspondence.html","id":"ref-usage","dir":"Reference","previous_headings":"","what":"Usage","title":"Generate togfen correspondence for two geographies — estimate_tongfen_single_correspondence","text":"","code":"estimate_tongfen_single_correspondence( geo1, geo2, geo1_uid, geo2_uid, tolerance = 1, computation_crs = NULL, robust = FALSE )"},{"path":"https://mountainmath.github.io/tongfen/reference/estimate_tongfen_single_correspondence.html","id":"arguments","dir":"Reference","previous_headings":"","what":"Arguments","title":"Generate togfen correspondence for two geographies — estimate_tongfen_single_correspondence","text":"geo1 input geometry 1 class sf geo2 input geometry 2 class sf geo1_uid (unique) identifier column geo1 geo2_uid (unique) identifier column geo2 tolerance tolerance (projected coordinate units) feature matching computation_crs optional crs computation carried , defaults crs geo1 robust boolean parameter, ensure geometries valid set TRUE","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/estimate_tongfen_single_correspondence.html","id":"value","dir":"Reference","previous_headings":"","what":"Value","title":"Generate togfen correspondence for two geographies — estimate_tongfen_single_correspondence","text":"correspondence table linking geo1_uid geo2_uid unique TongfenID TongfenUID columns enumerate common geometry.","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/get_correspondence_ca_census_for.html","id":null,"dir":"Reference","previous_headings":"","what":"Get StatCan DA or DB level correspondence file — get_correspondence_ca_census_for","title":"Get StatCan DA or DB level correspondence file — get_correspondence_ca_census_for","text":"Joins StatCan correspodence files several census years","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/get_correspondence_ca_census_for.html","id":"ref-usage","dir":"Reference","previous_headings":"","what":"Usage","title":"Get StatCan DA or DB level correspondence file — get_correspondence_ca_census_for","text":"","code":"get_correspondence_ca_census_for(years, level, refresh = FALSE)"},{"path":"https://mountainmath.github.io/tongfen/reference/get_correspondence_ca_census_for.html","id":"arguments","dir":"Reference","previous_headings":"","what":"Arguments","title":"Get StatCan DA or DB level correspondence file — get_correspondence_ca_census_for","text":"years list census years level geographic level, DA DB refresh reload correspondence files, default `FALSE`","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/get_correspondence_ca_census_for.html","id":"value","dir":"Reference","previous_headings":"","what":"Value","title":"Get StatCan DA or DB level correspondence file — get_correspondence_ca_census_for","text":"tibble correspondence table`spanning years","code":""},{"path":[]},{"path":"https://mountainmath.github.io/tongfen/reference/get_single_correspondence_ca_census_for.html","id":"ref-usage","dir":"Reference","previous_headings":"","what":"Usage","title":"Get StatCan DA or DB level correspondence file — get_single_correspondence_ca_census_for","text":"","code":"get_single_correspondence_ca_census_for( year, level = c(\"DA\", \"DB\"), refresh = FALSE )"},{"path":"https://mountainmath.github.io/tongfen/reference/get_single_correspondence_ca_census_for.html","id":"arguments","dir":"Reference","previous_headings":"","what":"Arguments","title":"Get StatCan DA or DB level correspondence file — get_single_correspondence_ca_census_for","text":"year census year, 2006 2021 supported level geographic level, DA DB refresh reload correspondence files, default `FALSE`","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/get_single_correspondence_ca_census_for.html","id":"value","dir":"Reference","previous_headings":"","what":"Value","title":"Get StatCan DA or DB level correspondence file — get_single_correspondence_ca_census_for","text":"tibble correspondence table`","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/get_tongfen_ca_census.html","id":null,"dir":"Reference","previous_headings":"","what":"Togfen data from several Canadian censuses — get_tongfen_ca_census","title":"Togfen data from several Canadian censuses — get_tongfen_ca_census","text":"Get data several Candian censuses common geography. Requires sf cancensus package available","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/get_tongfen_ca_census.html","id":"ref-usage","dir":"Reference","previous_headings":"","what":"Usage","title":"Togfen data from several Canadian censuses — get_tongfen_ca_census","text":"","code":"get_tongfen_ca_census( regions, meta, level = \"CT\", method = \"statcan\", base_geo = NULL, na.rm = FALSE, tolerance = 50, quiet = FALSE, refresh = FALSE, crs = NULL, data_transform = function(d) d )"},{"path":"https://mountainmath.github.io/tongfen/reference/get_tongfen_ca_census.html","id":"arguments","dir":"Reference","previous_headings":"","what":"Arguments","title":"Togfen data from several Canadian censuses — get_tongfen_ca_census","text":"regions census region list, inclusive list GeoUIDs across censuses meta metadata census veraiables aggregate, example returned meta_for_ca_census_vectors. level aggregation level return data (default \"CT\") method tongfen method, options \"statcan\" (default), \"estimate\", \"identifier\". * \"statcan\" method builds common geography using Statistics Canada correspondence files, point method works \"DB\", \"DA\" \"CT\" levels. * \"estimate\" uses `estimate_tongfen_correspondence` build common geography scratch based geographies. * \"identifier\" assumes regions identical geographic identifier identical, builds correspondence regions unmatched geographic identifiers. base_geo base census year build common geography , `NULL` (default) return geographic data na.rm logical, determines NA values treated aggregating variables, default `FALSE` tolerance tolerance `estimate_tongen_correspondence` metres, default value 50 metres, used method 'estimate' 'identifier' quiet suppress download progress output, default `FALSE` refresh optional character, refresh data cache call, (default `FALSE`) crs optional CRS transform data , use spatial intersections method 'identifier' 'estimate', defaults `3347` (Statistics Canada Lambert) intersections data_transform optional transform function applied census data returned cancensus","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/get_tongfen_ca_census.html","id":"value","dir":"Reference","previous_headings":"","what":"Value","title":"Togfen data from several Canadian censuses — get_tongfen_ca_census","text":"dataframe variables common geography","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/get_tongfen_ca_census.html","id":"ref-examples","dir":"Reference","previous_headings":"","what":"Examples","title":"Togfen data from several Canadian censuses — get_tongfen_ca_census","text":"","code":"# Get rent data for census years 2001 through 2016 if (FALSE) { # \\dontrun{ rent_variables <- c(rent_2001=\"v_CA01_1667\",rent_2016=\"v_CA16_4901\", rent_2011=\"v_CA11N_2292\",rent_2006=\"v_CA06_2050\") meta <- meta_for_ca_census_vectors(rent_variables) regions=list(CMA=\"59933\") rent_data <- get_tongfen_ca_census(regions=regions, meta=meta, quiet=TRUE, method=\"estimate\", level=\"CT\", base_geo = \"CA16\") } # }"},{"path":"https://mountainmath.github.io/tongfen/reference/get_tongfen_ca_census_ct_from_da.html","id":null,"dir":"Reference","previous_headings":"","what":"Canadian census CT level tongfen via DA correspondence — get_tongfen_ca_census_ct_from_da","title":"Canadian census CT level tongfen via DA correspondence — get_tongfen_ca_census_ct_from_da","text":"Grab variables several censuses common geography. Requires sf package available return CT level data","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/get_tongfen_ca_census_ct_from_da.html","id":"ref-usage","dir":"Reference","previous_headings":"","what":"Usage","title":"Canadian census CT level tongfen via DA correspondence — get_tongfen_ca_census_ct_from_da","text":"","code":"get_tongfen_ca_census_ct_from_da( regions, vectors, geo_format = NA, use_cache = TRUE, na.rm = TRUE, quiet = TRUE )"},{"path":"https://mountainmath.github.io/tongfen/reference/get_tongfen_ca_census_ct_from_da.html","id":"arguments","dir":"Reference","previous_headings":"","what":"Arguments","title":"Canadian census CT level tongfen via DA correspondence — get_tongfen_ca_census_ct_from_da","text":"regions census region list, inclusive list GeoUIDs across censuses vectors List cancensus vectors, can come different census years geo_format `NA` get variables 'sf' also get geographic data use_cache logical, passed `cancensus::get_census` regulate caching na.rm logical, determines NA values treated aggregating variables quiet suppress download progress output, default `TRUE`","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/get_tongfen_ca_census_ct_from_da.html","id":"value","dir":"Reference","previous_headings":"","what":"Value","title":"Canadian census CT level tongfen via DA correspondence — get_tongfen_ca_census_ct_from_da","text":"dataframe variables common geography","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/get_tongfen_census_ct.html","id":null,"dir":"Reference","previous_headings":"","what":"Canadian census CT level tongfen — get_tongfen_census_ct","title":"Canadian census CT level tongfen — get_tongfen_census_ct","text":"Grab variables several censuses common geography. Requires sf package available return CT level data","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/get_tongfen_census_ct.html","id":"ref-usage","dir":"Reference","previous_headings":"","what":"Usage","title":"Canadian census CT level tongfen — get_tongfen_census_ct","text":"","code":"get_tongfen_census_ct( regions, vectors, geo_format = NA, na.rm = TRUE, quiet = TRUE, refresh = FALSE )"},{"path":"https://mountainmath.github.io/tongfen/reference/get_tongfen_census_ct.html","id":"arguments","dir":"Reference","previous_headings":"","what":"Arguments","title":"Canadian census CT level tongfen — get_tongfen_census_ct","text":"regions census region list, inclusive list GeoUIDs across censuses vectors List cancensus vectors, can come different census years geo_format geographic format returned data, 'sf' sf format `NA“ na.rm remove NA values aggregating values, default `TRUE` quiet suppress download progress output, default `FALSE` refresh optional character, refresh data cache call","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/get_tongfen_census_ct.html","id":"value","dir":"Reference","previous_headings":"","what":"Value","title":"Canadian census CT level tongfen — get_tongfen_census_ct","text":"dataframe census variables common geography","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/get_tongfen_census_da.html","id":null,"dir":"Reference","previous_headings":"","what":"Canadian Census DA level tongfen — get_tongfen_census_da","title":"Canadian Census DA level tongfen — get_tongfen_census_da","text":"Grab variables several censuses common geography. Requires sf package available return CT level data","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/get_tongfen_census_da.html","id":"ref-usage","dir":"Reference","previous_headings":"","what":"Usage","title":"Canadian Census DA level tongfen — get_tongfen_census_da","text":"","code":"get_tongfen_census_da( regions, vectors, geo_format = NA, use_cache = TRUE, na.rm = TRUE, quiet = TRUE )"},{"path":"https://mountainmath.github.io/tongfen/reference/get_tongfen_census_da.html","id":"arguments","dir":"Reference","previous_headings":"","what":"Arguments","title":"Canadian Census DA level tongfen — get_tongfen_census_da","text":"regions census region list, inclusive list GeoUIDs across censuses vectors List cancensus vectors, can come different census years geo_format `NA` get variables 'sf' also get geographic data use_cache logical, passed `cancensus::get_census` regulate caching na.rm logical, determines NA values treated aggregating variables quiet suppress download progress output, default `TRUE`","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/get_tongfen_census_da.html","id":"value","dir":"Reference","previous_headings":"","what":"Value","title":"Canadian Census DA level tongfen — get_tongfen_census_da","text":"dataframe variables common geography","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/get_tongfen_correspondence_ca_census.html","id":null,"dir":"Reference","previous_headings":"","what":"Get StatCan correspondence data — get_tongfen_correspondence_ca_census","title":"Get StatCan correspondence data — get_tongfen_correspondence_ca_census","text":"Get correspondence file several Candian censuses common geography. Requires sf cancensus package available","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/get_tongfen_correspondence_ca_census.html","id":"ref-usage","dir":"Reference","previous_headings":"","what":"Usage","title":"Get StatCan correspondence data — get_tongfen_correspondence_ca_census","text":"","code":"get_tongfen_correspondence_ca_census( geo_datasets, regions, level = \"CT\", method = \"statcan\", tolerance = 50, quiet = FALSE, refresh = FALSE, crs = 3347 )"},{"path":"https://mountainmath.github.io/tongfen/reference/get_tongfen_correspondence_ca_census.html","id":"arguments","dir":"Reference","previous_headings":"","what":"Arguments","title":"Get StatCan correspondence data — get_tongfen_correspondence_ca_census","text":"geo_datasets vector census geography dataset identifiers regions census region list, inclusive list GeoUIDs across censuses level aggregation level return data (default \"CT\") method tongfen method, options \"statcan\" (default), \"estimate\", \"identifier\". * \"statcan\" method builds common geography using Statistics Canada correspondence files, point method works \"DB\", \"DA\" \"CT\" levels. * \"estimate\" uses `estimate_tongfen_correspondence` build common geography scratch based geographies. * \"identifier\" assumes regions identical geographic identifier identical, builds correspondence regions unmatched geographic identifiers. tolerance tolerance `estimate_tongen_correspondence` metres, default value 50 metres, used method 'estimate' 'identifier' quiet suppress download progress output, default `FALSE` refresh optional character, refresh data cache call, (default `FALSE`) crs CRS use spatial intersections method 'identifier' 'estimate', default `3347` (Statistics Canada Lambert)","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/get_tongfen_correspondence_ca_census.html","id":"value","dir":"Reference","previous_headings":"","what":"Value","title":"Get StatCan correspondence data — get_tongfen_correspondence_ca_census","text":"dataframe multi-census correspondence file","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/get_tongfen_correspondence_ca_census.html","id":"ref-examples","dir":"Reference","previous_headings":"","what":"Examples","title":"Get StatCan correspondence data — get_tongfen_correspondence_ca_census","text":"","code":"# Get correspondance files between CTs in 2006 and 2016 censuses in Vancouver CMA if (FALSE) { # \\dontrun{ correspondence <- get_tongfen_correspondence_ca_census(geo_datasets=c('CA06','CA16'), regions=list(CMA=\"59933\"),level='CT') } # }"},{"path":"https://mountainmath.github.io/tongfen/reference/get_tongfen_correspondence_us_census.html","id":null,"dir":"Reference","previous_headings":"","what":"Get correspondence table for US census geographies — get_tongfen_correspondence_us_census","title":"Get correspondence table for US census geographies — get_tongfen_correspondence_us_census","text":"Builds correspondence table matching US census geographies across censuses, based relationship files published US Census Bureau. Censuses requested sit two get traversed way, Census Bureau publishes relationship files consecutive censuses. relationship files geometric overlays list every sliver along boundaries shifted slightly. get cut via `min_area_share`, keeping chain unrelated regions one common geography. correspondence layer reaches back one census get_tongfen_us_census. 1990 census available `dec1990` , Census Bureau retired 1990 API endpoint, 1990 data brought means, example NHGIS via ipumsr package, handed tongfen_aggregate together correspondence table.","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/get_tongfen_correspondence_us_census.html","id":"ref-usage","dir":"Reference","previous_headings":"","what":"Usage","title":"Get correspondence table for US census geographies — get_tongfen_correspondence_us_census","text":"","code":"get_tongfen_correspondence_us_census( datasets, regions, level = \"tract\", min_area_share = 0.01, cache_path = getOption(\"tongfen.cache_path\") )"},{"path":"https://mountainmath.github.io/tongfen/reference/get_tongfen_correspondence_us_census.html","id":"arguments","dir":"Reference","previous_headings":"","what":"Arguments","title":"Get correspondence table for US census geographies — get_tongfen_correspondence_us_census","text":"datasets vector censuses match , valid values `dec1990`, `dec2000`, `dec2010` `dec2020` census tracts, `dec2000` `dec2020` county subdivisions. least two censuses needed. regions list regions query correspondence . stage, valid list vector states, .e. `regions = list(state=c(\"CA\",\"\"))` level aggregation level, stage valid levels 'tract' 'county subdivision'. min_area_share minimum share area two geographies common count related, default `0.01`. Census Bureau relationship files list every geometric overlap, lowering pulls slivers along boundaries shifted slightly chains unrelated regions one common geography. Raising gives finer common geographies risk separating regions change. region ever dropped, parts slivers largest part kept. cache_path optional path cache relationship files , defaults `tongfen.cache_path` option falls back temporary directory","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/get_tongfen_correspondence_us_census.html","id":"value","dir":"Reference","previous_headings":"","what":"Value","title":"Get correspondence table for US census geographies — get_tongfen_correspondence_us_census","text":"tibble one row per census geography, GEOID column requested census, common geography identified `TongfenID` `TongfenUID`.","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/get_tongfen_correspondence_us_census.html","id":"ref-examples","dir":"Reference","previous_headings":"","what":"Examples","title":"Get correspondence table for US census geographies — get_tongfen_correspondence_us_census","text":"","code":"# Match up census tracts for the 1990 and 2000 censuses in Rhode Island if (FALSE) { # \\dontrun{ correspondence <- get_tongfen_correspondence_us_census(datasets = c(\"dec1990\",\"dec2000\"), regions = list(state=\"RI\")) } # }"},{"path":"https://mountainmath.github.io/tongfen/reference/get_tongfen_us_census.html","id":null,"dir":"Reference","previous_headings":"","what":"Get US census data for 2000 and 2010 census on common census tract based geography — get_tongfen_us_census","title":"Get US census data for 2000 and 2010 census on common census tract based geography — get_tongfen_us_census","text":"wraps data acquisition via tidycensus package tongfen common geography single convenience function. Data available 2000, 2010 2020 censuses, Census Bureau retired 1990 API endpoint. tongfen 1990 data, obtain elsewhere combine correspondence table get_tongfen_correspondence_us_census via tongfen_aggregate.","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/get_tongfen_us_census.html","id":"ref-usage","dir":"Reference","previous_headings":"","what":"Usage","title":"Get US census data for 2000 and 2010 census on common census tract based geography — get_tongfen_us_census","text":"","code":"get_tongfen_us_census( regions, meta, level = \"tract\", survey = \"census\", base_geo = NULL, min_area_share = 0.01, sumfile = NULL )"},{"path":"https://mountainmath.github.io/tongfen/reference/get_tongfen_us_census.html","id":"arguments","dir":"Reference","previous_headings":"","what":"Arguments","title":"Get US census data for 2000 and 2010 census on common census tract based geography — get_tongfen_us_census","text":"regions list regions query data . stage, valid list vector states, .e. `regions = list(state=c(\"CA\",\"\"))“ meta metadata variables retrieve level aggregation level return data . stage, valid levels 'tract' 'county subdivision'. survey survey get data , supported options \"census\" base_geo census year use base geography, default `2010`. min_area_share minimum share area two geographies common count related, default `0.01`, see get_tongfen_correspondence_us_census. sumfile summary file read variables , either single value used censuses vector named dataset, example `c(dec2010=\"sf1\", dec2020=\"dhc\")`. Default `NULL`, leaves choice tidycensus. Note tidycensus defaults 2020 census PL 94-171 redistricting file, 2020 variables need `sumfile=\"dhc\"`.","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/get_tongfen_us_census.html","id":"value","dir":"Reference","previous_headings":"","what":"Value","title":"Get US census data for 2000 and 2010 census on common census tract based geography — get_tongfen_us_census","text":"sf object (wide form) census variables census year suffix (separated underdcore \"_\").","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/get_tongfen_us_census.html","id":"ref-examples","dir":"Reference","previous_headings":"","what":"Examples","title":"Get US census data for 2000 and 2010 census on common census tract based geography — get_tongfen_us_census","text":"","code":"# Get US census data on population and households for 2000 and 2010 censuses on a uniform geography # based on census tracts. if (FALSE) { # \\dontrun{ variables=c(population=\"H011001\",households=\"H013001\") meta <- c(2000,2010) %>% lapply(function(year){ v <- variables %>% setNames(paste0(names(.),\"_\",year)) meta_for_additive_variables(paste0(\"dec\",year),v) }) %>% bind_rows() census_data <- get_tongfen_us_census(regions = list(state=\"CA\"), meta=meta, level=\"tract\") %>% mutate(change=population_2010/households_2010-population_2000/households_2000) } # }"},{"path":"https://mountainmath.github.io/tongfen/reference/meta_for_additive_variables.html","id":null,"dir":"Reference","previous_headings":"","what":"Generate tongfen metadata for additive variables — meta_for_additive_variables","title":"Generate tongfen metadata for additive variables — meta_for_additive_variables","text":"Generates metadata used tongfen_aggregate. Variables need additive like counts.","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/meta_for_additive_variables.html","id":"ref-usage","dir":"Reference","previous_headings":"","what":"Usage","title":"Generate tongfen metadata for additive variables — meta_for_additive_variables","text":"","code":"meta_for_additive_variables(dataset, variables)"},{"path":"https://mountainmath.github.io/tongfen/reference/meta_for_additive_variables.html","id":"arguments","dir":"Reference","previous_headings":"","what":"Arguments","title":"Generate tongfen metadata for additive variables — meta_for_additive_variables","text":"dataset identifier dataset contianing variable variables (named) vecotor additive variables","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/meta_for_additive_variables.html","id":"value","dir":"Reference","previous_headings":"","what":"Value","title":"Generate tongfen metadata for additive variables — meta_for_additive_variables","text":"tibble used tongfen_aggregate","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/meta_for_additive_variables.html","id":"ref-examples","dir":"Reference","previous_headings":"","what":"Examples","title":"Generate tongfen metadata for additive variables — meta_for_additive_variables","text":"","code":"# Get metadata for additive variable Population for the CA16 and CA06 datasets if (FALSE) { # \\dontrun{ meta <- meta_for_additive_variables(c(\"CA06\",\"CA16\"),\"Population\") } # }"},{"path":"https://mountainmath.github.io/tongfen/reference/meta_for_ca_census_vectors.html","id":null,"dir":"Reference","previous_headings":"","what":"Generate metadata from Candian census vectors — meta_for_ca_census_vectors","title":"Generate metadata from Candian census vectors — meta_for_ca_census_vectors","text":"Build tibble information aggregate variables given vectors Queries list_census_variables obtain needed information add vectors needed aggregation","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/meta_for_ca_census_vectors.html","id":"ref-usage","dir":"Reference","previous_headings":"","what":"Usage","title":"Generate metadata from Candian census vectors — meta_for_ca_census_vectors","text":"","code":"meta_for_ca_census_vectors(vectors)"},{"path":"https://mountainmath.github.io/tongfen/reference/meta_for_ca_census_vectors.html","id":"arguments","dir":"Reference","previous_headings":"","what":"Arguments","title":"Generate metadata from Candian census vectors — meta_for_ca_census_vectors","text":"vectors list variables query","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/meta_for_ca_census_vectors.html","id":"value","dir":"Reference","previous_headings":"","what":"Value","title":"Generate metadata from Candian census vectors — meta_for_ca_census_vectors","text":"tidy dataframe metadata information requested variables additional variables needed tongfen operations","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/meta_for_ca_census_vectors.html","id":"ref-examples","dir":"Reference","previous_headings":"","what":"Examples","title":"Generate metadata from Candian census vectors — meta_for_ca_census_vectors","text":"","code":"# Build metadata for vectors if (FALSE) { # \\dontrun{ meta <- meta_for_ca_census_vectors(\"v_CA16_4836\",\"v_CA16_4838\",\"v_CA16_4899\") } # }"},{"path":"https://mountainmath.github.io/tongfen/reference/proportional_reaggregate.html","id":null,"dir":"Reference","previous_headings":"","what":"Dasymetric downsampling — proportional_reaggregate","title":"Dasymetric downsampling — proportional_reaggregate","text":"Proportionally re-aggregate hierarchical data lower-level w.r.t. values *base* variable Also handles cases lower level data may available blinded times filling data higher level Data lower aggregation levels may add accurate aggregate counts. function distributes aggregate level counts proportionally (population) containing lower level geographic regions.","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/proportional_reaggregate.html","id":"ref-usage","dir":"Reference","previous_headings":"","what":"Usage","title":"Dasymetric downsampling — proportional_reaggregate","text":"","code":"proportional_reaggregate( data, parent_data, geo_match, categories, base = \"Population\" )"},{"path":"https://mountainmath.github.io/tongfen/reference/proportional_reaggregate.html","id":"arguments","dir":"Reference","previous_headings":"","what":"Arguments","title":"Dasymetric downsampling — proportional_reaggregate","text":"data base geographic data parent_data Higher level geographic data geo_match named string informing column names match data parent_data categories Vector column names re-aggregate base Column name use proportional weighting re-aggregating, named vector column name category. Categries re-aggregated means set NA reaggregated base data NA values.","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/proportional_reaggregate.html","id":"value","dir":"Reference","previous_headings":"","what":"Value","title":"Dasymetric downsampling — proportional_reaggregate","text":"dataframe downsampled variables parent_data","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/proportional_reaggregate.html","id":"ref-examples","dir":"Reference","previous_headings":"","what":"Examples","title":"Dasymetric downsampling — proportional_reaggregate","text":"","code":"# Proportionally reaggregate visible minority data from dissemination area 2016 # census data to dissemination block geography, proportionally based on dissemination # block population if (FALSE) { # \\dontrun{ regions <- list(CSD=\"5915022\") variables <- cancensus::child_census_vectors(\"v_CA16_3954\") da_data <- cancensus::get_census(\"CA16\",regions=regions, vectors=setNames(variables$vector,variables$label), level=\"DA\") geo_data <- cancensus::get_census(\"CA16\",regions=regions,geo_format=\"sf\",level=\"DB\") db_data <- geo_data %>% proportional_reaggregate(da_data,c(\"DA_UID\"=\"GeoUID\"),variables$label) } # }"},{"path":"https://mountainmath.github.io/tongfen/reference/tongfen_aggregate.html","id":null,"dir":"Reference","previous_headings":"","what":"Perform tongfen according to correspondence — tongfen_aggregate","title":"Perform tongfen according to correspondence — tongfen_aggregate","text":"Aggregate variables secified meta several datasets according correspondence.","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/tongfen_aggregate.html","id":"ref-usage","dir":"Reference","previous_headings":"","what":"Usage","title":"Perform tongfen according to correspondence — tongfen_aggregate","text":"","code":"tongfen_aggregate( data, correspondence, meta = NULL, base_geo = NULL, na.rm = TRUE )"},{"path":"https://mountainmath.github.io/tongfen/reference/tongfen_aggregate.html","id":"arguments","dir":"Reference","previous_headings":"","what":"Arguments","title":"Perform tongfen according to correspondence — tongfen_aggregate","text":"data list datasets aggregated correspondence correspondence data gluing datasets meta metadata containing aggregation rules example returned `meta_for_ca_census_vectors` base_geo identifier data element base final geography , uses first data element `NULL` (default), expects `base_geo` element `names(data)`. na.rm logical, determines NA values treated aggregating variables, default `TRUE`","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/tongfen_aggregate.html","id":"value","dir":"Reference","previous_headings":"","what":"Value","title":"Perform tongfen according to correspondence — tongfen_aggregate","text":"aggregated dataset class sf base_geo NULL data type sf tibble otherwise.","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/tongfen_aggregate.html","id":"ref-examples","dir":"Reference","previous_headings":"","what":"Examples","title":"Perform tongfen according to correspondence — tongfen_aggregate","text":"","code":"# aggregate census tract level 2006 population data on common gepgraphy build through # correspondence from 2006 and 2016 census tracts in the City of Vancouver. if (FALSE) { # \\dontrun{ regions <- list(CSD=\"5915022\") geo1 <- cancensus::get_census(\"CA06\",regions=regions,geo_format='sf',level='CT') geo2 <- cancensus::get_census(\"CA16\",regions=regions,geo_format='sf',level='CT') meta <- meta_for_additive_variables(\"CA06\",\"Population\") correspondence <- get_tongfen_correspondence_ca_census(geo_datasets=c('CA06','CA16'), regions=regions,level='CT') result <- tongfen_aggregate(list(geo1 %>% rename(GeoUIDCA06=GeoUID), geo2 %>% rename(GeoUIDCA16=GeoUID)),correspondence,meta) } # }"},{"path":"https://mountainmath.github.io/tongfen/reference/tongfen_ca_census_ct.html","id":null,"dir":"Reference","previous_headings":"","what":"Canadian census CT level tongfen via identifier matching — tongfen_ca_census_ct","title":"Canadian census CT level tongfen via identifier matching — tongfen_ca_census_ct","text":"Aggregate variables common CTs, returns data2 new tiling matching data1 geography","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/tongfen_ca_census_ct.html","id":"ref-usage","dir":"Reference","previous_headings":"","what":"Usage","title":"Canadian census CT level tongfen via identifier matching — tongfen_ca_census_ct","text":"","code":"tongfen_ca_census_ct( data1, data2, data2_sum_vars, data2_group_vars = c(), na.rm = TRUE )"},{"path":"https://mountainmath.github.io/tongfen/reference/tongfen_ca_census_ct.html","id":"arguments","dir":"Reference","previous_headings":"","what":"Arguments","title":"Canadian census CT level tongfen via identifier matching — tongfen_ca_census_ct","text":"data1 cancensus CT level datatset year1 < year2 serve base common geography data2 cancensus CT level datatset year2 aggregated common geography data2_sum_vars vector variable names summed aggregating geographies data2_group_vars optional vector grouping variables na.rm optional parameter remove NA values summing, default = `TRUE`","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/tongfen_estimate.html","id":null,"dir":"Reference","previous_headings":"","what":"Estimate variable values for custom geography — tongfen_estimate","title":"Estimate variable values for custom geography — tongfen_estimate","text":"Estimates data source geometry onto target geometry using area-weighted interpolation. metadata specifies data aggregated, \"additive\" data like population counts summed proportionally area intersection, \"averages\" need additive \"parent\" count variables estimate weighted averages.","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/tongfen_estimate.html","id":"ref-usage","dir":"Reference","previous_headings":"","what":"Usage","title":"Estimate variable values for custom geography — tongfen_estimate","text":"","code":"tongfen_estimate(target, source, meta, na.rm = FALSE)"},{"path":"https://mountainmath.github.io/tongfen/reference/tongfen_estimate.html","id":"arguments","dir":"Reference","previous_headings":"","what":"Arguments","title":"Estimate variable values for custom geography — tongfen_estimate","text":"target custom geography estimate values source input geography values meta metadata variable aggregation, see `meta_for_additive_variables` `meta_for_ca_census_vectors` information construct metadata. na.rm remove NA values aggregating, default FALSE","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/tongfen_estimate.html","id":"value","dir":"Reference","previous_headings":"","what":"Value","title":"Estimate variable values for custom geography — tongfen_estimate","text":"`target` estimated quantities `source` specified `meta`","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/tongfen_estimate.html","id":"ref-examples","dir":"Reference","previous_headings":"","what":"Examples","title":"Estimate variable values for custom geography — tongfen_estimate","text":"","code":"# Estimate 2006 Population in the City of Vancouver dissemination ares on 2016 census geographies if (FALSE) { # \\dontrun{ geo1 <- cancensus::get_census(\"CA06\",regions=list(CSD=\"5915022\"),geo_format='sf',level='DA') geo2 <- cancensus::get_census(\"CA16\",regions=list(CSD=\"5915022\"),geo_format='sf',level='DA') meta <- meta_for_additive_variables(\"CA06\",\"Population\") result <- tongfen_estimate(geo2 %>% rename(Population_2016=Population),geo1,meta) } # }"},{"path":"https://mountainmath.github.io/tongfen/reference/tongfen_estimate_ca_census.html","id":null,"dir":"Reference","previous_headings":"","what":"Tongfen estimate data for given geometry — tongfen_estimate_ca_census","title":"Tongfen estimate data for given geometry — tongfen_estimate_ca_census","text":"Estimates values given census vectors given geometry using data specified level range. wrapper around `cancensus::get_intersecting_geometries` `tongfen_estimate`, optionally downsampling via `proportional_reaggregate`, streamline estimating Canadian census data custom geographies.","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/tongfen_estimate_ca_census.html","id":"ref-usage","dir":"Reference","previous_headings":"","what":"Usage","title":"Tongfen estimate data for given geometry — tongfen_estimate_ca_census","text":"","code":"tongfen_estimate_ca_census( geometry, meta, level, intersection_level = level, downsample_level = NULL, na.rm = FALSE, quiet = FALSE )"},{"path":"https://mountainmath.github.io/tongfen/reference/tongfen_estimate_ca_census.html","id":"arguments","dir":"Reference","previous_headings":"","what":"Arguments","title":"Tongfen estimate data for given geometry — tongfen_estimate_ca_census","text":"geometry geometry meta metadata census variables aggregate, example returned `meta_for_ca_census_vectors`. point function accepts variables census geography year. expand also allow estimates across multiple census geography years, requires attention detail. recommended apply due caution running function separately across several census geography years purpose comparing data across time naive application can lead systematic biases. level level use tongfen intersection_level level use geometry intersection, different tongfen level meta_for_ca_census_vectors. can set higher aggregation level conserve API points `get_intersecting_geometries` call. downsample_level default `NULL`, can geographic level lower `level`, case data downsamples geography level proportionally using value `downsample` column (must supplied) `meta` argument intersecting geometries. can lead accurate results. point allowed variables `downsample` column `meta` \"Population\", \"Households\" \"Dwellings\", can one variables. na.rm deal NA values, default FALSE. quiet suppress progress messages","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/tongfen_estimate_ca_census.html","id":"ref-examples","dir":"Reference","previous_headings":"","what":"Examples","title":"Tongfen estimate data for given geometry — tongfen_estimate_ca_census","text":"","code":"# Estimate a common geography for 2006 and 2016 dissemination areas in the City of Vancouver # based on the geographic data and check estimation errors if (FALSE) { # \\dontrun{ toronto_city_hall <- sf::st_point(c(-79.3839,43.6534)) %>% sf::st_sfc(crs=4326) %>% sf::st_transform(3348) %>% sf::st_buffer(1000) %>% sf::st_sf() meta <- meta_for_additive_variables(\"CA16\",\"Population\") data <- tongfen_estimate_ca_census(toronto_city_hall,meta,level=\"DA\",intersection_level=\"CT\") print(paste0(\"Approximately \",scales::comma(data$Population,accuracy=100), \" people live within a 1 km radius of Toronto City.\")) } # }"},{"path":"https://mountainmath.github.io/tongfen/reference/tongfen_tag_largest_overlap.html","id":null,"dir":"Reference","previous_headings":"","what":"Tag regions by largest overlap — tongfen_tag_largest_overlap","title":"Tag regions by largest overlap — tongfen_tag_largest_overlap","text":"tags regions `source` `target_id` region `target` largest overlap","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/tongfen_tag_largest_overlap.html","id":"ref-usage","dir":"Reference","previous_headings":"","what":"Usage","title":"Tag regions by largest overlap — tongfen_tag_largest_overlap","text":"","code":"tongfen_tag_largest_overlap(source, target, target_id)"},{"path":"https://mountainmath.github.io/tongfen/reference/tongfen_tag_largest_overlap.html","id":"arguments","dir":"Reference","previous_headings":"","what":"Arguments","title":"Tag regions by largest overlap — tongfen_tag_largest_overlap","text":"source input geography target custom geography target_id name column `target` table unique id (character)","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/tongfen_tag_largest_overlap.html","id":"value","dir":"Reference","previous_headings":"","what":"Value","title":"Tag regions by largest overlap — tongfen_tag_largest_overlap","text":"`source` extra column name `\"target_id\"` column `...overlap_fraction` proportion overlap target geometry respective `target_id`","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/tongfen_tag_largest_overlap.html","id":"ref-examples","dir":"Reference","previous_headings":"","what":"Examples","title":"Tag regions by largest overlap — tongfen_tag_largest_overlap","text":"","code":"# Estimate 2006 Populatino in the City of Vancouver dissemination ares on 2016 census geoographies if (FALSE) { # \\dontrun{ geo1 <- cancensus::get_census(\"CA06\",regions=list(CSD=\"5915022\"),geo_format='sf',level='DA') geo2 <- cancensus::get_census(\"CA16\",regions=list(CSD=\"5915022\"),geo_format='sf',level='DA') meta <- meta_for_additive_variables(\"CA06\",\"Population\") result <- tongfen_estimate(geo2 %>% rename(Population_2016=Population),geo1,meta) } # }"},{"path":"https://mountainmath.github.io/tongfen/reference/vancouver_elections_data_2015.html","id":null,"dir":"Reference","previous_headings":"","what":"A dataset with polling station votes data from the 2015 federal election in the Vancouver area — vancouver_elections_data_2015","title":"A dataset with polling station votes data from the 2015 federal election in the Vancouver area — vancouver_elections_data_2015","text":"dataset polling station votes data 2015 federal election Vancouver area","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/vancouver_elections_data_2015.html","id":"references","dir":"Reference","previous_headings":"","what":"References","title":"A dataset with polling station votes data from the 2015 federal election in the Vancouver area — vancouver_elections_data_2015","text":"https://www.elections.ca/content.aspx?section=res&dir=rep/&document=index&lang=e#42GE","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/vancouver_elections_data_2015.html","id":"author","dir":"Reference","previous_headings":"","what":"Author","title":"A dataset with polling station votes data from the 2015 federal election in the Vancouver area — vancouver_elections_data_2015","text":"Elections Canada","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/vancouver_elections_data_2019.html","id":null,"dir":"Reference","previous_headings":"","what":"A dataset with polling station votes data from the 2019 federal election in the Vancouver area — vancouver_elections_data_2019","title":"A dataset with polling station votes data from the 2019 federal election in the Vancouver area — vancouver_elections_data_2019","text":"dataset polling station votes data 2019 federal election Vancouver area","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/vancouver_elections_data_2019.html","id":"references","dir":"Reference","previous_headings":"","what":"References","title":"A dataset with polling station votes data from the 2019 federal election in the Vancouver area — vancouver_elections_data_2019","text":"https://www.elections.ca/content.aspx?section=res&dir=rep/&document=index&lang=e#43GE","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/vancouver_elections_data_2019.html","id":"author","dir":"Reference","previous_headings":"","what":"Author","title":"A dataset with polling station votes data from the 2019 federal election in the Vancouver area — vancouver_elections_data_2019","text":"Elections Canada","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/vancouver_elections_geos_2015.html","id":null,"dir":"Reference","previous_headings":"","what":"A dataset with polling district geographies from the 2015 federal election in the Vancouver area — vancouver_elections_geos_2015","title":"A dataset with polling district geographies from the 2015 federal election in the Vancouver area — vancouver_elections_geos_2015","text":"dataset polling district geographies 2015 federal election Vancouver area","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/vancouver_elections_geos_2015.html","id":"references","dir":"Reference","previous_headings":"","what":"References","title":"A dataset with polling district geographies from the 2015 federal election in the Vancouver area — vancouver_elections_geos_2015","text":"https://www.elections.ca/content.aspx?section=res&dir=rep/&document=index&lang=e#42GE","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/vancouver_elections_geos_2015.html","id":"author","dir":"Reference","previous_headings":"","what":"Author","title":"A dataset with polling district geographies from the 2015 federal election in the Vancouver area — vancouver_elections_geos_2015","text":"Elections Canada","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/vancouver_elections_geos_2019.html","id":null,"dir":"Reference","previous_headings":"","what":"A dataset with polling district geographies from the 2019 federal election in the Vancouver area — vancouver_elections_geos_2019","title":"A dataset with polling district geographies from the 2019 federal election in the Vancouver area — vancouver_elections_geos_2019","text":"dataset polling district geographies 2019 federal election Vancouver area","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/vancouver_elections_geos_2019.html","id":"references","dir":"Reference","previous_headings":"","what":"References","title":"A dataset with polling district geographies from the 2019 federal election in the Vancouver area — vancouver_elections_geos_2019","text":"https://www.elections.ca/content.aspx?section=res&dir=rep/&document=index&lang=e#43GE","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/vancouver_elections_geos_2019.html","id":"author","dir":"Reference","previous_headings":"","what":"Author","title":"A dataset with polling district geographies from the 2019 federal election in the Vancouver area — vancouver_elections_geos_2019","text":"Elections Canada","code":""},{"path":"https://mountainmath.github.io/tongfen/news/index.html","id":"tongfen-032","dir":"Changelog","previous_headings":"","what":"tongfen 0.3.2","title":"tongfen 0.3.2","text":"Fix compatibility issue changes {sf} package reliable GitHub action CRAN checks","code":""},{"path":[]},{"path":"https://mountainmath.github.io/tongfen/news/index.html","id":"major-changes-0-3-2","dir":"Changelog","previous_headings":"","what":"Major changes","title":"tongfen 0.3.2","text":"Added tongfen_estimate_ca_census function new CensusMapper endpoint, tying new {cancensus} functionality. ## Minor changes Custom impelementation tongfen_etimate finer control Fix compatibility issue changes {sf} package","code":""},{"path":"https://mountainmath.github.io/tongfen/news/index.html","id":"tongfen-03","dir":"Changelog","previous_headings":"","what":"tongfen 0.3","title":"tongfen 0.3","text":"CRAN release: 2020-11-04","code":""},{"path":"https://mountainmath.github.io/tongfen/news/index.html","id":"major-changes-0-3","dir":"Changelog","previous_headings":"","what":"Major changes","title":"tongfen 0.3","text":"Initial release","code":""}]
+[{"path":"https://mountainmath.github.io/tongfen/LICENSE.html","id":null,"dir":"","previous_headings":"","what":"MIT License","title":"MIT License","text":"Copyright (c) 2020 Jens von Bergmann Permission hereby granted, free charge, person obtaining copy software associated documentation files (“Software”), deal Software without restriction, including without limitation rights use, copy, modify, merge, publish, distribute, sublicense, /sell copies Software, permit persons Software furnished , subject following conditions: copyright notice permission notice shall included copies substantial portions Software. SOFTWARE PROVIDED “”, WITHOUT WARRANTY KIND, EXPRESS IMPLIED, INCLUDING LIMITED WARRANTIES MERCHANTABILITY, FITNESS PARTICULAR PURPOSE NONINFRINGEMENT. EVENT SHALL AUTHORS COPYRIGHT HOLDERS LIABLE CLAIM, DAMAGES LIABILITY, WHETHER ACTION CONTRACT, TORT OTHERWISE, ARISING , CONNECTION SOFTWARE USE DEALINGS SOFTWARE.","code":""},{"path":"https://mountainmath.github.io/tongfen/articles/tongfen.html","id":"population-change","dir":"Articles","previous_headings":"","what":"Population change","title":"General TongFen","text":"’s time go back original goal mapping population change. need specify aggregate population data, simply adding . meta_for_additive_variables convenience function generates appropriate metadata specifies deal data. ’s left add population data. choose 2001 base year clipped boundaries look better.","code":"meta <- meta_for_additive_variables(years,\"Population\") meta #> # A tibble: 4 × 8 #> variable dataset label type aggregation rule geo_dataset parent #> #> 1 Population 2001 Population_2001 Manual Additive Addi… 2001 NA #> 2 Population 2006 Population_2006 Manual Additive Addi… 2006 NA #> 3 Population 2011 Population_2011 Manual Additive Addi… 2011 NA #> 4 Population 2016 Population_2016 Manual Additive Addi… 2016 NA breaks = c(-0.15,-0.1,-0.075,-0.05,-0.025,0,0.025,0.05,0.1,0.2,0.3) labels = c(\"-15% to -10%\",\"-10% to -7.5%\",\"-7.5% to -5%\",\"-5% to -2.5%\",\"-2.5% to 0%\",\"0% to 2.5%\",\"2.5% to 5%\",\"5% to 10%\",\"10% to 20%\",\"20% to 30%\") colors <- RColorBrewer::brewer.pal(10,\"PiYG\") compute_population_change_metrics <- function(data) { geometric_average <- function(x,n){sign(x) * (exp(log(1+abs(x))/n)-1)} data %>% mutate(`2001 - 2006`=geometric_average((`Population_2006`-`Population_2001`)/`Population_2001`,5), `2006 - 2011`=geometric_average((`Population_2011`-`Population_2006`)/`Population_2006`,5), `2011 - 2016`=geometric_average((`Population_2016`-`Population_2011`)/`Population_2011`,5), `2001 - 2016`=geometric_average((`Population_2016`-`Population_2001`)/`Population_2001`,15)) %>% gather(key=\"Period\",value=\"Population Change\",c(\"2001 - 2006\",\"2006 - 2011\",\"2011 - 2016\",\"2001 - 2016\")) %>% mutate(Period=factor(Period,levels=c(\"2001 - 2006\",\"2006 - 2011\",\"2011 - 2016\",\"2001 - 2016\"))) %>% mutate(c=cut(`Population Change`,breaks=breaks, labels=labels)) } plot_data <- tongfen_aggregate(data,correspondence,meta=meta,base_geo = \"2001\") %>% compute_population_change_metrics() ggplot(plot_data,aes(fill=c)) + geom_sf(size=0.1) + scale_fill_manual(values=setNames(colors,labels)) + facet_wrap(\"Period\",ncol=2) + coord_sf(datum=NA) + labs(fill=\"Average Annual\\nPopulation Change\", title=\"Vancouver population change\", caption = \"StatCan Census 2001-2016\")"},{"path":"https://mountainmath.github.io/tongfen/articles/tongfen_anomalies.html","id":"population-timelines-for-toronto","dir":"Articles","previous_headings":"","what":"Population timelines for Toronto","title":"Geocoding anomalies in TongFen timelines","text":"example take population 1971 2011 censuses Statistics Canada tabulated 2016 dissemination areas, together 2016 population. data comes geography, need TongFen, data earlier years geocoded road network block face time, always match 2016 dissemination areas. Dissemination areas without population given year come back missing values. Changes missing value never considered surprising, set zero mark areas nobody got counted. area around Crescent Town shows problem looks like. population jumps back forth two areas, sum two fairly steady 1981 . People move back forth, homes got geocoded different dissemination area different years.","code":"years <- c(1971,seq(1981,2011,5)) vectors <- c(setNames(paste0(\"v_CA\",years,\"x16_1\"),years),\"2016\"=\"v_CA16_1\") timeline <- names(vectors) toronto <- get_census(\"CA16CT\",regions=list(CSD=\"3520005\"),vectors=vectors, level=\"DA\",geo_format=\"sf\",quiet=TRUE) %>% select(GeoUID,all_of(timeline)) %>% mutate(across(all_of(timeline),\\(x) coalesce(x,0))) crescent_town <- c(\"35204370\",\"35204765\") plot_timelines <- function(data) { data %>% st_drop_geometry() %>% pivot_longer(all_of(timeline),names_to=\"Year\",values_to=\"Population\") %>% ggplot(aes(x=Year,y=Population,colour=GeoUID,group=GeoUID)) + geom_line() + geom_point() + scale_y_continuous(labels=scales::comma,limits=c(0,NA)) } toronto %>% filter(GeoUID %in% crescent_town) %>% plot_timelines() + labs(title=\"Population in two neighbouring dissemination areas\")"},{"path":"https://mountainmath.github.io/tongfen/articles/tongfen_anomalies.html","id":"detecting-anomalies","dir":"Articles","previous_headings":"","what":"Detecting anomalies","title":"Geocoding anomalies in TongFen timelines","text":"tongfen_detect_anomalies lists regions surprising drops. decreases surprising, decrease needs large relative absolute terms. candidate regions finds neighbouring region takes away surprise joined, checks reduction large enough justify joining . regions identified GeoUID instead TongfenID function looks default. areas candidates, one neighbour best explains surprising drops . candidates find neighbour pair . Population drop real, example site gets cleared redevelopment, regions left alone.","code":"anomalies <- tongfen_detect_anomalies(toronto,timeline,id=\"GeoUID\",total_surprise_cutoff=0.4) anomalies %>% filter(GeoUID %in% crescent_town) #> # A tibble: 2 × 7 #> GeoUID surprise_count surprise_total period neighbour surprise_total_joined #> #> 1 35204370 3 0.933 1991-1… 35204765 0.616 #> 2 35204765 2 1.03 1981-1… 35204370 0.142 #> # ℹ 1 more variable: join anomalies %>% count(join) #> # A tibble: 2 × 2 #> join n #> #> 1 FALSE 254 #> 2 TRUE 288"},{"path":"https://mountainmath.github.io/tongfen/articles/tongfen_anomalies.html","id":"joining-regions","dir":"Articles","previous_headings":"","what":"Joining regions","title":"Geocoding anomalies in TongFen timelines","text":"tongfen_anomaly_joins joins regions qualify looks , joined regions can surprising drops complemented another neighbour. repeats regions left join. result lists regions got joined together identifier joined region now part round first got joined. tongfen_join_regions applies joins data, aggregating variables geometries regions get joined leaving others . map shows regions got joined around Crescent Town.","code":"joins <- tongfen_anomaly_joins(toronto,timeline,id=\"GeoUID\",total_surprise_cutoff=0.4) joins %>% filter(GeoUID %in% crescent_town) #> # A tibble: 2 × 3 #> GeoUID GeoUID_joined round #> #> 1 35204370 35204370 1 #> 2 35204765 35204370 1 toronto_joined <- tongfen_join_regions(toronto,joins,id=\"GeoUID\") c(original=nrow(toronto),joined=nrow(toronto_joined)) #> original joined #> 3702 3430 toronto_joined %>% filter(GeoUID %in% crescent_town) %>% plot_timelines() + labs(title=\"Population in the joined region\") bbox <- toronto %>% filter(GeoUID %in% crescent_town) %>% st_buffer(1500) %>% st_bbox() ggplot(toronto_joined %>% mutate(joined=GeoUID %in% joins$GeoUID_joined)) + geom_sf(aes(fill=joined),linewidth=0.1) + geom_sf(data=toronto,fill=NA,linewidth=0.1,linetype=\"dotted\") + scale_fill_manual(values=c(\"TRUE\"=\"steelblue\",\"FALSE\"=\"whitesmoke\"),guide=\"none\") + coord_sf(datum=NA,xlim=bbox[c(\"xmin\",\"xmax\")],ylim=bbox[c(\"ymin\",\"ymax\")]) + labs(title=\"Joined regions around Crescent Town\", caption=\"Joined regions in blue, original dissemination areas dotted\")"},{"path":"https://mountainmath.github.io/tongfen/articles/tongfen_anomalies.html","id":"tuning","dir":"Articles","previous_headings":"","what":"Tuning","title":"Geocoding anomalies in TongFen timelines","text":"Joining regions trades geographic detail consistency time, best make trade depends data application. parameters documented tongfen_detect_anomalies, important ones rel_scale abs_scale, relative absolute decrease change half way fully surprising. defaults 25% drop drop 200 tuned population counts regions size dissemination areas. total_surprise_cutoff, surprising timeline region needs become candidate. default 0.75 conservative, used 0.4 also pick less pronounced cases. cutoff_fact, surprise_reduction_const sum_fact determine much surprise neighbour needs take away regions get joined. Neighbours default determined intersecting geometries regions. can miss neighbours geometries simplified, case neighbours argument takes table identifiers neighbouring regions neighbours list spdep package.","code":"c(0.4,0.6,0.75) %>% lapply(\\(cutoff) tibble(total_surprise_cutoff=cutoff, regions_joined=tongfen_anomaly_joins(toronto,timeline,id=\"GeoUID\", total_surprise_cutoff=cutoff) %>% nrow())) %>% bind_rows() #> # A tibble: 3 × 2 #> total_surprise_cutoff regions_joined #> #> 1 0.4 477 #> 2 0.6 273 #> 3 0.75 152"},{"path":"https://mountainmath.github.io/tongfen/articles/tongfen_anomalies.html","id":"anomalies-in-tongfen-data","dir":"Articles","previous_headings":"","what":"Anomalies in TongFen data","title":"Geocoding anomalies in TongFen timelines","text":"functions work way data common geography built TongFen, regions identified TongfenID. example look dissemination area level population City Vancouver 2001 2021 censuses. Passing metadata tongfen_join_regions makes sure variables get aggregated right way, numeric variables part metadata assumed additive. TongfenUID joined regions lists dissemination areas made . Variables additive, like averages, can aggregated way variable averaged part data. alternative always works join regions correspondence common geography built , use joined correspondence aggregate original data. also way use joins data one used detect anomalies, example get average rents corrected geography.","code":"regions <- list(CSD=\"5915022\") datasets <- c(\"CA01\",\"CA06\",\"CA11\",\"CA16\",\"CA21\") meta <- meta_for_additive_variables(datasets,\"Population\") vancouver <- get_tongfen_ca_census(regions=regions,meta=meta,level=\"DA\",base_geo=\"CA21\",quiet=TRUE) joins <- tongfen_anomaly_joins(vancouver,paste0(\"Population_\",datasets),total_surprise_cutoff=0.4) joins #> # A tibble: 4 × 3 #> TongfenID TongfenID_joined round #> #> 1 59150762 59150762 1 #> 2 59153181 59150762 1 #> 3 59150765 59150765 1 #> 4 59150770 59150765 1 vancouver_joined <- tongfen_join_regions(vancouver,joins,meta) vancouver_joined %>% st_drop_geometry() %>% filter(TongfenID %in% joins$TongfenID_joined) %>% select(TongfenID,TongfenUID,starts_with(\"Population\")) #> # A tibble: 2 × 7 #> TongfenID TongfenUID Population_CA21 Population_CA01 Population_CA06 #> #> 1 59150762 GeoUIDCA01:59150762… 1412 1148 1266 #> 2 59150765 GeoUIDCA01:59150765… 4082 1895 2499 #> # ℹ 2 more variables: Population_CA11 , Population_CA16 correspondence <- get_tongfen_correspondence_ca_census(geo_datasets=datasets,regions=regions, level=\"DA\",quiet=TRUE) %>% tongfen_join_correspondence(joins) rent_meta <- meta_for_ca_census_vectors(c(rent_2006=\"v_CA06_2050\",rent_2016=\"v_CA16_4901\")) rent_data <- c(\"CA06\",\"CA16\") %>% lapply(\\(ds) get_census(ds,regions=regions,level=\"DA\",labels=\"short\",quiet=TRUE, vectors=rent_meta %>% filter(geo_dataset==ds) %>% pull(variable), geo_format=if (ds==\"CA16\") \"sf\" else NA) %>% rename(!!paste0(\"GeoUID\",ds):=\"GeoUID\")) %>% setNames(c(\"CA06\",\"CA16\")) rents <- tongfen_aggregate(rent_data,correspondence,rent_meta,base_geo=\"CA16\") rents %>% st_drop_geometry() %>% filter(TongfenID %in% joins$TongfenID_joined) %>% select(TongfenID,rent_2006,rent_2016) #> # A tibble: 2 × 3 #> TongfenID rent_2006 rent_2016 #> #> 1 59150762 452. 674. #> 2 59150765 474. 913."},{"path":"https://mountainmath.github.io/tongfen/articles/tongfen_ca.html","id":"aggregating-up-data-across-regions","dir":"Articles","previous_headings":"","what":"Aggregating up data across regions","title":"TongFen for Canadian census data","text":"Another application simply aggregating variables selection regions single census. Suppose want understand share renters Vancouver School District, well share renter households spending 30% income housing. 53% Vancouver School District households rent, 45% shelter cost burdened.","code":"vectors <- c(\"v_CA16_4836\",\"v_CA16_4838\",\"v_CA16_4899\") meta=meta_for_ca_census_vectors(vectors) %>% bind_rows(meta_for_additive_variables(\"CA16\",c(\"Population\",\"Dwellings\",\"Households\"))) vsb_regions <- list(CSD=c(\"5915022\",\"5915803\"), CT=c(\"9330069.01\",\"9330069.02\",\"9330069.00\")) vsb <- get_census(\"CA16\",regions=vsb_regions,vectors=meta$variable,labels=\"short\") vsb <- aggregate_data_with_meta(vsb, meta) %>% mutate(Total=v_CA16_4836,Renters=v_CA16_4838,rent_poor=v_CA16_4899/100) %>% mutate(rent_share=Renters/Total)"},{"path":"https://mountainmath.github.io/tongfen/articles/tongfen_ca.html","id":"change-in-vancouver-children","dir":"Articles","previous_headings":"Aggregating up data across regions","what":"Change in Vancouver children","title":"TongFen for Canadian census data","text":"Data can also obtained dissemination area level. example look change children aged 0 14 Vancouver School District.","code":"variables <- c(\"2016_0-14\"=\"v_CA16_4\", \"2011_0-4\"=\"v_CA11F_8\",\"2011_5-9\"=\"v_CA11F_11\",\"2011_10-14\"=\"v_CA11F_14\") meta <- meta_for_ca_census_vectors(variables) %>% bind_rows(meta_for_additive_variables(c(\"CA11\",\"CA16\"),\"Population\")) children_data <- get_tongfen_ca_census(regions = vsb_regions, meta = meta, level=\"DA\", base_geo = \"CA16\", quiet = TRUE) %>% mutate(`2011_0-14`=purrr::reduce(select(sf::st_set_geometry(.,NULL), starts_with(\"2011_\")), `+`)) %>% mutate(change=`2016_0-14`/Population_CA16-`2011_0-14`/Population_CA11) ggplot(children_data,aes(fill=change)) + geom_sf(size=0.1) + scale_fill_gradient2(labels=scales::percent) + coord_sf(datum=NA) + labs(title=\"Percentage point change in share of children aged 0-14 between 2011 and 2016\",fill=NULL)"},{"path":"https://mountainmath.github.io/tongfen/articles/tongfen_us.html","id":"bridging-to-the-2020-census","dir":"Articles","previous_headings":"","what":"Bridging to the 2020 census","title":"TongFen for US census data","text":"works across 2010 2020 censuses. Two things change. 2020 census renamed variables, population occupied housing units H8_001N households H3_002N. live Demographic Housing Characteristics file, whereas tidycensus reads PL 94-171 redistricting file 2020 default, point right one via sumfile. takes single value censuses, one named dataset .","code":"meta_2020 <- bind_rows( meta_for_additive_variables(\"dec2010\",c(population_2010=\"H011001\", households_2010=\"H013001\")), meta_for_additive_variables(\"dec2020\",c(population_2020=\"H8_001N\", households_2020=\"H3_002N\"))) census_data_2020 <- get_tongfen_us_census(regions = list(state=\"CA\"), meta=meta_2020, level=\"tract\", sumfile=c(dec2020=\"dhc\")) %>% mutate(change=population_2020/households_2020-population_2010/households_2010) census_data_2020 %>% mutate(c=cut(change,c(-Inf,-0.5,-0.3,-0.2,-0.1,0,0.1,0.2,0.3,0.5,Inf))) %>% ggplot() + geom_sf(aes(fill=c), size=0.05) + scale_fill_brewer(palette = \"RdYlGn\") + labs(title=\"Bay area change in average household size 2010-2020\", fill=NULL) + coord_sf(datum=NA,xlim=c(-122.6,-121.7),ylim=c(37.2,37.9))"},{"path":"https://mountainmath.github.io/tongfen/articles/tongfen_us.html","id":"notes-on-the-common-geography","dir":"Articles","previous_headings":"","what":"Notes on the common geography","title":"TongFen for US census data","text":"Census Bureau relationship files correspondences built geometric overlays, list every place two censuses’ geographies intersect, including slivers along boundaries shifted metres. Chaining together merges regions nothing , min_area_share sets much area two regions common count related. default 0.01 works well, raising gives finer common geographies risk separating regions genuinely change. region ever dropped, region’s parts fall cutoff largest part kept. Correspondence tables reach back one census data . Census Bureau retired 1990 API endpoint, tidycensus fetch 1990 data, get_tongfen_correspondence_us_census match 1990 tracts later censuses. use , get 1990 data elsewhere, example NHGIS via ipumsr package, hand tongfen_aggregate along correspondence table.","code":""},{"path":"https://mountainmath.github.io/tongfen/authors.html","id":null,"dir":"","previous_headings":"","what":"Authors","title":"Authors and Citation","text":"Jens von Bergmann. Author, maintainer. creator maintainer","code":""},{"path":"https://mountainmath.github.io/tongfen/authors.html","id":"citation","dir":"","previous_headings":"","what":"Citation","title":"Authors and Citation","text":"von Bergmann J (2026). tongfen: Make Data Based Different Geographies Comparable. R package version 0.3.9, https://github.com/mountainMath/tongfen.","code":"@Manual{, title = {tongfen: Make Data Based on Different Geographies Comparable}, author = {Jens {von Bergmann}}, year = {2026}, note = {R package version 0.3.9}, url = {https://github.com/mountainMath/tongfen}, }"},{"path":"https://mountainmath.github.io/tongfen/index.html","id":"tongfen","dir":"","previous_headings":"","what":"Make Data Based on Different Geographies Comparable","title":"Make Data Based on Different Geographies Comparable","text":"TongFen (通分) means convert two fractions least common denominator, typically preparation manipulation like addition subtraction. English, ’s mouthful sounds complicated. Chinese word , TongFen, makes process appear simple. working geospatial datasets often want compare data given different regions. example census data election data. data two different censuses. properly compare data first need convert common geography. process quite analogous process TongFen fractions, appropriate term give simple name. Using tongfen package, preparing data disparate geographies comparison converting common geography easy typing tongfen.","code":""},{"path":"https://mountainmath.github.io/tongfen/index.html","id":"reference","dir":"","previous_headings":"","what":"Reference","title":"Make Data Based on Different Geographies Comparable","text":"TongFen home page reference guide","code":""},{"path":"https://mountainmath.github.io/tongfen/index.html","id":"installing-the-package","dir":"","previous_headings":"","what":"Installing the package","title":"Make Data Based on Different Geographies Comparable","text":"latest development version can installed GitHub.","code":"install.packages(\"tongfen\") remotes::install_github(\"mountainmath/tongfen\") library(tongfen)"},{"path":"https://mountainmath.github.io/tongfen/index.html","id":"caching-correspondence-files","dir":"","previous_headings":"","what":"Caching correspondence files","title":"Make Data Based on Different Geographies Comparable","text":"get_tongfen_ca_census get_tongfen_correspondence_ca_census methods make use StatCan correspondence files run method = \"statcan\". Statistics Canada longer allows programmatic downloads files, package downloads mirror hosting files parquet format. speed process useful permanently cache files instead download every session. caching desired, set either options(\"tongfen.cache_path\"=\"\") Sys.setenv(\"tongfen.cache_path\"=\"\") options(\"custom_data_path\"=\"\") .Rprofile .Renviron file. Cached files checked mirror per session get downloaded changed, offline cached files used .","code":""},{"path":"https://mountainmath.github.io/tongfen/index.html","id":"general-tongfen","dir":"","previous_headings":"","what":"General TongFen","title":"Make Data Based on Different Geographies Comparable","text":"tongfen package build around following basic TongFen workflow: Given list datasets diverse geographies, generate correspondence table links geographies specifies aggregate (least) common geography via estimate_tongfen_correspondence. generate metadata specifies variables can aggregated , meta_for_additive_variables function additive variables. Use correspondence table metadata generate dataset variables original datasets aggregated common geography via tongfen_aggregate. convenience function validate geographic TongFen fit via area comparison available via check_tongfen_areas, allows explore deal spatial mismatches TongFen.","code":""},{"path":"https://mountainmath.github.io/tongfen/index.html","id":"aggregation-of-variables","dir":"","previous_headings":"General TongFen","what":"Aggregation of variables","title":"Make Data Based on Different Geographies Comparable","text":"Finding common tiling several different yet congruent geographies one part problem TongFen addresses, aggregating variables part. tongfen package deals using metadata table specifies variables aggregated. ’s simplest form values simply added . meta_for_additive_variables convenience function builds metadata additive variables. Metadata non-additive variables like averages, ratios percentages needs care build, requires additional information parent variable specifies denominator average, ratio percentage. data, like medians, can’t aggregated , although tongfen can provide estimates medians aggregated geographies treating averages.","code":""},{"path":"https://mountainmath.github.io/tongfen/index.html","id":"packaged-data","dir":"","previous_headings":"General TongFen","what":"Packaged data","title":"Make Data Based on Different Geographies Comparable","text":"package ships subset voting data Elections Canada 42nd 43rd federal elections well polling district geographies 42nd 43rd. facilitates running example vignette polling districts without download external data. available open data covered Open Government Licence - Canada.","code":""},{"path":"https://mountainmath.github.io/tongfen/index.html","id":"data-specific-implementations","dir":"","previous_headings":"","what":"Data-specific implementations","title":"Make Data Based on Different Geographies Comparable","text":"need TongFen comes frequently certain types geographies. Census geographies one example. cases data sources come correspondence files go beyond geographic matchup also join regions alleviate data integrity problems like geocoding issues. cases can worthwhile wrap data acquisition TongFen one convenience function, also extend TongFen method parameter allow external correspondence files used.","code":""},{"path":"https://mountainmath.github.io/tongfen/index.html","id":"canadian-census-data","dir":"","previous_headings":"Data-specific implementations","what":"Canadian census data","title":"Make Data Based on Different Geographies Comparable","text":"package well-integrated work Canadian census data two essential ways. * meta_for_ca_census_vectors builds rich metadata given list Canadian census variables utilizing metadata available via CensusMapper. particular, automates proper aggregation non-count variables like averages, ratios percentages. * get_tongfen_ca_census wraps process data acquisition (via CensusMapper cancensus package tongfen one convenience function. time adds TongFen method = \"statcan\" option uses Statistics Canada correspondence files build common geography. * get_tongfen_correspondence_ca_census function breaks correspondence generation aid process accessing Statistics Canada correspondence files (better integration generating correspondences Canadian census geographies general) facilitate mixing non-census data coming census geographies, like example CMHC data.","code":""},{"path":"https://mountainmath.github.io/tongfen/index.html","id":"us-census-data","dir":"","previous_headings":"Data-specific implementations","what":"US census data","title":"Make Data Based on Different Geographies Comparable","text":"get_tongfen_us_census integrates data acquisition (via tidycensus package) TongFen, using US Census Bureau relationship files build common geography. get_tongfen_correspondence_us_census breaks correspondence generation US Census Bureau relationship files, tongfen data comes census geographies obtained means. relationship files cached us_data folder cache path described .","code":""},{"path":"https://mountainmath.github.io/tongfen/index.html","id":"other-implementations","dir":"","previous_headings":"","what":"Other implementations","title":"Make Data Based on Different Geographies Comparable","text":"tongfen package open add extensions specialized data sources, well extensions existing ones.","code":""},{"path":"https://mountainmath.github.io/tongfen/index.html","id":"fixed-target-geography-estimation","dir":"","previous_headings":"","what":"Fixed target geography estimation","title":"Make Data Based on Different Geographies Comparable","text":"geographies aren’t sufficiently congruent target geography fixed, won’t able use tongfen methods compute data common geography instead rely estimates. tongfen_estimate makes assumption underlying geographies returns estimates data target geography. uses area-weighted interpolation achieve , can refined dasymetric estimates using proportional_reaggregate function. method example works independent nature underlying geographies, comes heavy price estimate. useful research purposes also need methods estimate errors introduces effects subsequent analysis results. Methods facilitate still active development.","code":""},{"path":"https://mountainmath.github.io/tongfen/index.html","id":"cite-tongfen","dir":"","previous_headings":"Fixed target geography estimation","what":"Cite tongfen","title":"Make Data Based on Different Geographies Comparable","text":"wish cite tongfen: von Bergmann, J. (2026). tongfen: R package Make Data Based Different Geographies Comparable. v0.3.9. DOI: 10.32614/CRAN.package.tongfen BibTeX entry LaTeX users ","code":"@Manual{tongfen, author = {Jens {von Bergmann}}, title = {tongfen: R package to Make Data Based on Different Geographies Comparable}, year = {2026}, doi = {10.32614/CRAN.package.tongfen}, note = {R package version 0.3.9}, url = {https://mountainmath.github.io/tongfen/}, }"},{"path":"https://mountainmath.github.io/tongfen/reference/add_census_ca_base_variables.html","id":null,"dir":"Reference","previous_headings":"","what":"Generate metadata from Canadian census vectors — add_census_ca_base_variables","title":"Generate metadata from Canadian census vectors — add_census_ca_base_variables","text":"Add Population, Dwellings, Household counts metadata","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/add_census_ca_base_variables.html","id":"ref-usage","dir":"Reference","previous_headings":"","what":"Usage","title":"Generate metadata from Canadian census vectors — add_census_ca_base_variables","text":"","code":"add_census_ca_base_variables(meta)"},{"path":"https://mountainmath.github.io/tongfen/reference/add_census_ca_base_variables.html","id":"arguments","dir":"Reference","previous_headings":"","what":"Arguments","title":"Generate metadata from Canadian census vectors — add_census_ca_base_variables","text":"meta tibble metadata example provided `meta_for_ca_census_vectors`","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/add_census_ca_base_variables.html","id":"value","dir":"Reference","previous_headings":"","what":"Value","title":"Generate metadata from Canadian census vectors — add_census_ca_base_variables","text":"tibble metadata","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/aggregate_data_with_meta.html","id":null,"dir":"Reference","previous_headings":"","what":"Aggregate variables in grouped data — aggregate_data_with_meta","title":"Aggregate variables in grouped data — aggregate_data_with_meta","text":"Aggregate census data , assumes data grouped aggregation Uses data meta determine aggregate ","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/aggregate_data_with_meta.html","id":"ref-usage","dir":"Reference","previous_headings":"","what":"Usage","title":"Aggregate variables in grouped data — aggregate_data_with_meta","text":"","code":"aggregate_data_with_meta(data, meta, geo = FALSE, na.rm = TRUE, quiet = FALSE)"},{"path":"https://mountainmath.github.io/tongfen/reference/aggregate_data_with_meta.html","id":"arguments","dir":"Reference","previous_headings":"","what":"Arguments","title":"Aggregate variables in grouped data — aggregate_data_with_meta","text":"data census data obtained get_census call, grouped TongfenID meta list variables aggregation information obtained meta_for_vectors geo logical, also aggregate geographic data na.rm logical, NA values ignored carried . quiet logical, emit messages set `TRUE`","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/aggregate_data_with_meta.html","id":"value","dir":"Reference","previous_headings":"","what":"Value","title":"Aggregate variables in grouped data — aggregate_data_with_meta","text":"data frame variables aggregated new common geography","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/aggregate_data_with_meta.html","id":"ref-examples","dir":"Reference","previous_headings":"","what":"Examples","title":"Aggregate variables in grouped data — aggregate_data_with_meta","text":"","code":"# Aggregate population from DA level to grouped by CT_UID if (FALSE) { # \\dontrun{ geo <- cancensus::get_census(\"CA06\",regions=list(CSD=\"5915022\"),level='DA') meta <- meta_for_additive_variables(\"CA06\",\"Population\") result <- aggregate_data_with_meta(geo %>% group_by(CT_UID),meta) } # }"},{"path":"https://mountainmath.github.io/tongfen/reference/check_tongfen_areas.html","id":null,"dir":"Reference","previous_headings":"","what":"Check geographic integrity — check_tongfen_areas","title":"Check geographic integrity — check_tongfen_areas","text":"Sanity check areas estimated tongfen correspondence. useful example total extent geo1 geo2 differ regions edges large difference overlap. result diagnostic, pass/fail test. Geographies different years simplified independently differ water features cut , sizable area mismatch mean regions matched incorrectly.","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/check_tongfen_areas.html","id":"ref-usage","dir":"Reference","previous_headings":"","what":"Usage","title":"Check geographic integrity — check_tongfen_areas","text":"","code":"check_tongfen_areas(data, correspondence)"},{"path":"https://mountainmath.github.io/tongfen/reference/check_tongfen_areas.html","id":"arguments","dir":"Reference","previous_headings":"","what":"Arguments","title":"Check geographic integrity — check_tongfen_areas","text":"data list geographic data class sf correspondence Correspondence table columns unique geographic identifiers geographies TongfenID (optionally TongfenUID TongfenMethod) returned `estimate_tongfen_correspondence`.","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/check_tongfen_areas.html","id":"value","dir":"Reference","previous_headings":"","what":"Value","title":"Check geographic integrity — check_tongfen_areas","text":"table columns `TongfenID`, geo_identifiers, areas aggregated regions corresponding geographic identifier column, tongfen estimation method maximum log ratio areas.","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/check_tongfen_areas.html","id":"ref-examples","dir":"Reference","previous_headings":"","what":"Examples","title":"Check geographic integrity — check_tongfen_areas","text":"","code":"# Estimate a common geography for 2006 and 2016 dissemination areas in the City of Vancouver # based on the geographic data and check estimation errors if (FALSE) { # \\dontrun{ regions <- list(CSD=\"5915022\") data_06 <- cancensus::get_census(\"CA06\",regions=regions,geo_format='sf',level=\"DA\") %>% rename(GeoUID_06=GeoUID) data_16 <- cancensus::get_census(\"CA16\",regions=regions,geo_format=\"sf\",level=\"DA\") %>% rename(GeoUID_16=GeoUID) correspondence <- estimate_tongfen_correspondence(list(data_06, data_16), c(\"GeoUID_06\",\"GeoUID_16\")) area_check <- check_tongfen_areas(list(data_06, data_16),correspondence) } # }"},{"path":"https://mountainmath.github.io/tongfen/reference/check_tongfen_single_areas.html","id":null,"dir":"Reference","previous_headings":"","what":"Check geographic integrity — check_tongfen_single_areas","title":"Check geographic integrity — check_tongfen_single_areas","text":"Sanity check areas estimated tongfen correspondence. useful example total extent geo1 geo2 differ regions edges large difference overlap.","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/check_tongfen_single_areas.html","id":"ref-usage","dir":"Reference","previous_headings":"","what":"Usage","title":"Check geographic integrity — check_tongfen_single_areas","text":"","code":"check_tongfen_single_areas(geo1, geo2, correspondence)"},{"path":"https://mountainmath.github.io/tongfen/reference/check_tongfen_single_areas.html","id":"arguments","dir":"Reference","previous_headings":"","what":"Arguments","title":"Check geographic integrity — check_tongfen_single_areas","text":"geo1 input geometry 1 class sf geo2 input geometry 2 class sf correspondence Correspondence table `geo1` `geo2` e.g. returned `estimate_tongfen_correspondence`.","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/check_tongfen_single_areas.html","id":"value","dir":"Reference","previous_headings":"","what":"Value","title":"Check geographic integrity — check_tongfen_single_areas","text":"table columns `TongfenID`, `area1` `area2`, row corresponds unique `TongfenID` `correspondence` table columns hold areas regions aggregated `geo1` `geo2`.`","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/estimate_tongfen_correspondence.html","id":null,"dir":"Reference","previous_headings":"","what":"Generate tongfen correspondence for list of geographies — estimate_tongfen_correspondence","title":"Generate tongfen correspondence for list of geographies — estimate_tongfen_correspondence","text":"Get correspondence data arbitrary congruent geometries. Congruent means one can obtain common tiling aggregating several sub-geometries two input geo data. Worst case scenario common tiling given unioning sub-geometries finer common tiling.","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/estimate_tongfen_correspondence.html","id":"ref-usage","dir":"Reference","previous_headings":"","what":"Usage","title":"Generate tongfen correspondence for list of geographies — estimate_tongfen_correspondence","text":"","code":"estimate_tongfen_correspondence( data, geo_identifiers, method = \"estimate\", tolerance = 50, computation_crs = NULL )"},{"path":"https://mountainmath.github.io/tongfen/reference/estimate_tongfen_correspondence.html","id":"arguments","dir":"Reference","previous_headings":"","what":"Arguments","title":"Generate tongfen correspondence for list of geographies — estimate_tongfen_correspondence","text":"data list geometries class sf geo_identifiers vector unique geographic identifiers list entry data. method aggregation method. Possible values \"estimate\" \"identifier\". \"estimate\" estimates correspondence purely geographic data. \"identifier\" assumes regions identical geo_identifiers , uses \"estimate\" method remaining regions. Default \"estimate\". tolerance tolerance (projected coordinate units `computation_crs`) feature matching computation_crs optional crs computation carried , defaults crs first entry data parameter.","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/estimate_tongfen_correspondence.html","id":"value","dir":"Reference","previous_headings":"","what":"Value","title":"Generate tongfen correspondence for list of geographies — estimate_tongfen_correspondence","text":"correspondence table linking geo1_uid geo2_uid unique TongfenID TongfenUID columns enumerate common geometry.","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/estimate_tongfen_correspondence.html","id":"ref-examples","dir":"Reference","previous_headings":"","what":"Examples","title":"Generate tongfen correspondence for list of geographies — estimate_tongfen_correspondence","text":"","code":"# Estimate a common geography for 2006 and 2016 dissemination areas in the City of Vancouver # based on the geographic data. if (FALSE) { # \\dontrun{ regions <- list(CSD=\"5915022\") data_06 <- cancensus::get_census(\"CA06\",regions=regions,geo_format='sf',level=\"DA\") %>% rename(GeoUID_06=GeoUID) data_16 <- cancensus::get_census(\"CA16\",regions=regions,geo_format=\"sf\",level=\"DA\") %>% rename(GeoUID_16=GeoUID) correspondence <- estimate_tongfen_correspondence(list(data_06, data_16), c(\"GeoUID_06\",\"GeoUID_16\")) } # }"},{"path":"https://mountainmath.github.io/tongfen/reference/estimate_tongfen_single_correspondence.html","id":null,"dir":"Reference","previous_headings":"","what":"Generate tongfen correspondence for two geographies — estimate_tongfen_single_correspondence","title":"Generate tongfen correspondence for two geographies — estimate_tongfen_single_correspondence","text":"Get correspondence data arbitrary congruent geometries. Congruent means one can obtain common tiling aggregating several sub-geometries two input geo data. Worst case scenario common tiling given unioning sub-geometries finer common tiling.","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/estimate_tongfen_single_correspondence.html","id":"ref-usage","dir":"Reference","previous_headings":"","what":"Usage","title":"Generate tongfen correspondence for two geographies — estimate_tongfen_single_correspondence","text":"","code":"estimate_tongfen_single_correspondence( geo1, geo2, geo1_uid, geo2_uid, tolerance = 1, computation_crs = NULL, robust = FALSE )"},{"path":"https://mountainmath.github.io/tongfen/reference/estimate_tongfen_single_correspondence.html","id":"arguments","dir":"Reference","previous_headings":"","what":"Arguments","title":"Generate tongfen correspondence for two geographies — estimate_tongfen_single_correspondence","text":"geo1 input geometry 1 class sf geo2 input geometry 2 class sf geo1_uid (unique) identifier column geo1 geo2_uid (unique) identifier column geo2 tolerance tolerance (projected coordinate units) feature matching computation_crs optional crs computation carried , defaults crs geo1 robust boolean parameter, ensure geometries valid set TRUE","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/estimate_tongfen_single_correspondence.html","id":"value","dir":"Reference","previous_headings":"","what":"Value","title":"Generate tongfen correspondence for two geographies — estimate_tongfen_single_correspondence","text":"correspondence table linking geo1_uid geo2_uid unique TongfenID TongfenUID columns enumerate common geometry.","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/get_correspondence_ca_census_for.html","id":null,"dir":"Reference","previous_headings":"","what":"Get StatCan DA or DB level correspondence file — get_correspondence_ca_census_for","title":"Get StatCan DA or DB level correspondence file — get_correspondence_ca_census_for","text":"Joins StatCan correspondence files several census years","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/get_correspondence_ca_census_for.html","id":"ref-usage","dir":"Reference","previous_headings":"","what":"Usage","title":"Get StatCan DA or DB level correspondence file — get_correspondence_ca_census_for","text":"","code":"get_correspondence_ca_census_for(years, level, refresh = FALSE)"},{"path":"https://mountainmath.github.io/tongfen/reference/get_correspondence_ca_census_for.html","id":"arguments","dir":"Reference","previous_headings":"","what":"Arguments","title":"Get StatCan DA or DB level correspondence file — get_correspondence_ca_census_for","text":"years list census years level geographic level, DA DB refresh reload correspondence files, default `FALSE`","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/get_correspondence_ca_census_for.html","id":"value","dir":"Reference","previous_headings":"","what":"Value","title":"Get StatCan DA or DB level correspondence file — get_correspondence_ca_census_for","text":"tibble correspondence table`spanning years","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/get_single_correspondence_ca_census_for.html","id":null,"dir":"Reference","previous_headings":"","what":"Get StatCan DA or DB level correspondence file — get_single_correspondence_ca_census_for","title":"Get StatCan DA or DB level correspondence file — get_single_correspondence_ca_census_for","text":"correspondence files downloaded mirror Statistics Canada correspondence files cached tongfen cache directory. cached files checked mirror per session get downloaded changed. location mirror can changed via `tongfen.statcan_correspondence_url` option.","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/get_single_correspondence_ca_census_for.html","id":"ref-usage","dir":"Reference","previous_headings":"","what":"Usage","title":"Get StatCan DA or DB level correspondence file — get_single_correspondence_ca_census_for","text":"","code":"get_single_correspondence_ca_census_for( year, level = c(\"DA\", \"DB\"), refresh = FALSE )"},{"path":"https://mountainmath.github.io/tongfen/reference/get_single_correspondence_ca_census_for.html","id":"arguments","dir":"Reference","previous_headings":"","what":"Arguments","title":"Get StatCan DA or DB level correspondence file — get_single_correspondence_ca_census_for","text":"year census year, 2006 2021 supported level geographic level, DA DB refresh reload correspondence files, default `FALSE`","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/get_single_correspondence_ca_census_for.html","id":"value","dir":"Reference","previous_headings":"","what":"Value","title":"Get StatCan DA or DB level correspondence file — get_single_correspondence_ca_census_for","text":"tibble correspondence table`","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/get_tongfen_ca_census.html","id":null,"dir":"Reference","previous_headings":"","what":"Tongfen data from several Canadian censuses — get_tongfen_ca_census","title":"Tongfen data from several Canadian censuses — get_tongfen_ca_census","text":"Get data several Canadian censuses common geography. Requires sf cancensus package available","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/get_tongfen_ca_census.html","id":"ref-usage","dir":"Reference","previous_headings":"","what":"Usage","title":"Tongfen data from several Canadian censuses — get_tongfen_ca_census","text":"","code":"get_tongfen_ca_census( regions, meta, level = \"CT\", method = \"statcan\", base_geo = NULL, na.rm = FALSE, tolerance = 50, quiet = FALSE, refresh = FALSE, crs = NULL, data_transform = function(d) d )"},{"path":"https://mountainmath.github.io/tongfen/reference/get_tongfen_ca_census.html","id":"arguments","dir":"Reference","previous_headings":"","what":"Arguments","title":"Tongfen data from several Canadian censuses — get_tongfen_ca_census","text":"regions census region list, inclusive list GeoUIDs across censuses meta metadata census variables aggregate, example returned meta_for_ca_census_vectors. level aggregation level return data (default \"CT\") method tongfen method, options \"statcan\" (default), \"estimate\", \"identifier\". * \"statcan\" method builds common geography using Statistics Canada correspondence files, point method works \"DB\", \"DA\" \"CT\" levels. * \"estimate\" uses `estimate_tongfen_correspondence` build common geography scratch based geographies. * \"identifier\" assumes regions identical geographic identifier identical, builds correspondence regions unmatched geographic identifiers. base_geo base census year build common geography , `NULL` (default) return geographic data na.rm logical, determines NA values treated aggregating variables, default `FALSE` tolerance tolerance `estimate_tongfen_correspondence` metres, default value 50 metres, used method 'estimate' 'identifier' quiet suppress download progress output, default `FALSE` refresh optional character, refresh data cache call, (default `FALSE`) crs optional CRS transform data , use spatial intersections method 'identifier' 'estimate', defaults `3347` (Statistics Canada Lambert) intersections data_transform optional transform function applied census data returned cancensus","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/get_tongfen_ca_census.html","id":"value","dir":"Reference","previous_headings":"","what":"Value","title":"Tongfen data from several Canadian censuses — get_tongfen_ca_census","text":"dataframe variables common geography","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/get_tongfen_ca_census.html","id":"ref-examples","dir":"Reference","previous_headings":"","what":"Examples","title":"Tongfen data from several Canadian censuses — get_tongfen_ca_census","text":"","code":"# Get rent data for census years 2001 through 2016 if (FALSE) { # \\dontrun{ rent_variables <- c(rent_2001=\"v_CA01_1667\",rent_2016=\"v_CA16_4901\", rent_2011=\"v_CA11N_2292\",rent_2006=\"v_CA06_2050\") meta <- meta_for_ca_census_vectors(rent_variables) regions=list(CMA=\"59933\") rent_data <- get_tongfen_ca_census(regions=regions, meta=meta, quiet=TRUE, method=\"estimate\", level=\"CT\", base_geo = \"CA16\") } # }"},{"path":"https://mountainmath.github.io/tongfen/reference/get_tongfen_ca_census_ct_from_da.html","id":null,"dir":"Reference","previous_headings":"","what":"Canadian census CT level tongfen via DA correspondence — get_tongfen_ca_census_ct_from_da","title":"Canadian census CT level tongfen via DA correspondence — get_tongfen_ca_census_ct_from_da","text":"Grab variables several censuses common geography. Requires sf package available return CT level data","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/get_tongfen_ca_census_ct_from_da.html","id":"ref-usage","dir":"Reference","previous_headings":"","what":"Usage","title":"Canadian census CT level tongfen via DA correspondence — get_tongfen_ca_census_ct_from_da","text":"","code":"get_tongfen_ca_census_ct_from_da( regions, vectors, geo_format = NA, use_cache = TRUE, na.rm = TRUE, quiet = TRUE )"},{"path":"https://mountainmath.github.io/tongfen/reference/get_tongfen_ca_census_ct_from_da.html","id":"arguments","dir":"Reference","previous_headings":"","what":"Arguments","title":"Canadian census CT level tongfen via DA correspondence — get_tongfen_ca_census_ct_from_da","text":"regions census region list, inclusive list GeoUIDs across censuses vectors List cancensus vectors, can come different census years geo_format `NA` get variables 'sf' also get geographic data use_cache logical, passed `cancensus::get_census` regulate caching na.rm logical, determines NA values treated aggregating variables quiet suppress download progress output, default `TRUE`","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/get_tongfen_ca_census_ct_from_da.html","id":"value","dir":"Reference","previous_headings":"","what":"Value","title":"Canadian census CT level tongfen via DA correspondence — get_tongfen_ca_census_ct_from_da","text":"dataframe variables common geography","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/get_tongfen_census_ct.html","id":null,"dir":"Reference","previous_headings":"","what":"Canadian census CT level tongfen — get_tongfen_census_ct","title":"Canadian census CT level tongfen — get_tongfen_census_ct","text":"Grab variables several censuses common geography. Requires sf package available return CT level data","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/get_tongfen_census_ct.html","id":"ref-usage","dir":"Reference","previous_headings":"","what":"Usage","title":"Canadian census CT level tongfen — get_tongfen_census_ct","text":"","code":"get_tongfen_census_ct( regions, vectors, geo_format = NA, na.rm = TRUE, quiet = TRUE, refresh = FALSE )"},{"path":"https://mountainmath.github.io/tongfen/reference/get_tongfen_census_ct.html","id":"arguments","dir":"Reference","previous_headings":"","what":"Arguments","title":"Canadian census CT level tongfen — get_tongfen_census_ct","text":"regions census region list, inclusive list GeoUIDs across censuses vectors List cancensus vectors, can come different census years geo_format geographic format returned data, 'sf' sf format `NA“ na.rm remove NA values aggregating values, default `TRUE` quiet suppress download progress output, default `FALSE` refresh optional character, refresh data cache call","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/get_tongfen_census_ct.html","id":"value","dir":"Reference","previous_headings":"","what":"Value","title":"Canadian census CT level tongfen — get_tongfen_census_ct","text":"dataframe census variables common geography","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/get_tongfen_census_da.html","id":null,"dir":"Reference","previous_headings":"","what":"Canadian Census DA level tongfen — get_tongfen_census_da","title":"Canadian Census DA level tongfen — get_tongfen_census_da","text":"Grab variables several censuses common geography. Requires sf package available return CT level data","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/get_tongfen_census_da.html","id":"ref-usage","dir":"Reference","previous_headings":"","what":"Usage","title":"Canadian Census DA level tongfen — get_tongfen_census_da","text":"","code":"get_tongfen_census_da( regions, vectors, geo_format = NA, use_cache = TRUE, na.rm = TRUE, quiet = TRUE )"},{"path":"https://mountainmath.github.io/tongfen/reference/get_tongfen_census_da.html","id":"arguments","dir":"Reference","previous_headings":"","what":"Arguments","title":"Canadian Census DA level tongfen — get_tongfen_census_da","text":"regions census region list, inclusive list GeoUIDs across censuses vectors List cancensus vectors, can come different census years geo_format `NA` get variables 'sf' also get geographic data use_cache logical, passed `cancensus::get_census` regulate caching na.rm logical, determines NA values treated aggregating variables quiet suppress download progress output, default `TRUE`","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/get_tongfen_census_da.html","id":"value","dir":"Reference","previous_headings":"","what":"Value","title":"Canadian Census DA level tongfen — get_tongfen_census_da","text":"dataframe variables common geography","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/get_tongfen_correspondence_ca_census.html","id":null,"dir":"Reference","previous_headings":"","what":"Get StatCan correspondence data — get_tongfen_correspondence_ca_census","title":"Get StatCan correspondence data — get_tongfen_correspondence_ca_census","text":"Get correspondence file several Canadian censuses common geography. Requires sf cancensus package available","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/get_tongfen_correspondence_ca_census.html","id":"ref-usage","dir":"Reference","previous_headings":"","what":"Usage","title":"Get StatCan correspondence data — get_tongfen_correspondence_ca_census","text":"","code":"get_tongfen_correspondence_ca_census( geo_datasets, regions, level = \"CT\", method = \"statcan\", tolerance = 50, quiet = FALSE, refresh = FALSE, crs = 3347 )"},{"path":"https://mountainmath.github.io/tongfen/reference/get_tongfen_correspondence_ca_census.html","id":"arguments","dir":"Reference","previous_headings":"","what":"Arguments","title":"Get StatCan correspondence data — get_tongfen_correspondence_ca_census","text":"geo_datasets vector census geography dataset identifiers regions census region list, inclusive list GeoUIDs across censuses level aggregation level return data (default \"CT\") method tongfen method, options \"statcan\" (default), \"estimate\", \"identifier\". * \"statcan\" method builds common geography using Statistics Canada correspondence files, point method works \"DB\", \"DA\" \"CT\" levels. * \"estimate\" uses `estimate_tongfen_correspondence` build common geography scratch based geographies. * \"identifier\" assumes regions identical geographic identifier identical, builds correspondence regions unmatched geographic identifiers. tolerance tolerance `estimate_tongfen_correspondence` metres, default value 50 metres, used method 'estimate' 'identifier' quiet suppress download progress output, default `FALSE` refresh optional character, refresh data cache call, (default `FALSE`) crs CRS use spatial intersections method 'identifier' 'estimate', default `3347` (Statistics Canada Lambert)","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/get_tongfen_correspondence_ca_census.html","id":"value","dir":"Reference","previous_headings":"","what":"Value","title":"Get StatCan correspondence data — get_tongfen_correspondence_ca_census","text":"dataframe multi-census correspondence file","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/get_tongfen_correspondence_ca_census.html","id":"ref-examples","dir":"Reference","previous_headings":"","what":"Examples","title":"Get StatCan correspondence data — get_tongfen_correspondence_ca_census","text":"","code":"# Get correspondance files between CTs in 2006 and 2016 censuses in Vancouver CMA if (FALSE) { # \\dontrun{ correspondence <- get_tongfen_correspondence_ca_census(geo_datasets=c('CA06','CA16'), regions=list(CMA=\"59933\"),level='CT') } # }"},{"path":"https://mountainmath.github.io/tongfen/reference/get_tongfen_correspondence_us_census.html","id":null,"dir":"Reference","previous_headings":"","what":"Get correspondence table for US census geographies — get_tongfen_correspondence_us_census","title":"Get correspondence table for US census geographies — get_tongfen_correspondence_us_census","text":"Builds correspondence table matching US census geographies across censuses, based relationship files published US Census Bureau. Censuses requested sit two get traversed way, Census Bureau publishes relationship files consecutive censuses. relationship files geometric overlays list every sliver along boundaries shifted slightly. get cut via `min_area_share`, keeping chain unrelated regions one common geography. correspondence layer reaches back one census get_tongfen_us_census. 1990 census available `dec1990` , Census Bureau retired 1990 API endpoint, 1990 data brought means, example NHGIS via ipumsr package, handed tongfen_aggregate together correspondence table.","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/get_tongfen_correspondence_us_census.html","id":"ref-usage","dir":"Reference","previous_headings":"","what":"Usage","title":"Get correspondence table for US census geographies — get_tongfen_correspondence_us_census","text":"","code":"get_tongfen_correspondence_us_census( datasets, regions, level = \"tract\", min_area_share = 0.01, cache_path = getOption(\"tongfen.cache_path\") )"},{"path":"https://mountainmath.github.io/tongfen/reference/get_tongfen_correspondence_us_census.html","id":"arguments","dir":"Reference","previous_headings":"","what":"Arguments","title":"Get correspondence table for US census geographies — get_tongfen_correspondence_us_census","text":"datasets vector censuses match , valid values `dec1990`, `dec2000`, `dec2010` `dec2020` census tracts, `dec2000` `dec2020` county subdivisions. least two censuses needed. regions list regions query correspondence . stage, valid list vector states, .e. `regions = list(state=c(\"CA\",\"\"))` level aggregation level, stage valid levels 'tract' 'county subdivision'. min_area_share minimum share area two geographies common count related, default `0.01`. Census Bureau relationship files list every geometric overlap, lowering pulls slivers along boundaries shifted slightly chains unrelated regions one common geography. Raising gives finer common geographies risk separating regions change. region ever dropped, parts slivers largest part kept. cache_path optional path cache relationship files , defaults `tongfen.cache_path` option. set `tongfen.cache_path` environment variable `custom_data_path` option used, falling back temporary directory","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/get_tongfen_correspondence_us_census.html","id":"value","dir":"Reference","previous_headings":"","what":"Value","title":"Get correspondence table for US census geographies — get_tongfen_correspondence_us_census","text":"tibble one row per census geography, GEOID column requested census, common geography identified `TongfenID` `TongfenUID`.","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/get_tongfen_correspondence_us_census.html","id":"ref-examples","dir":"Reference","previous_headings":"","what":"Examples","title":"Get correspondence table for US census geographies — get_tongfen_correspondence_us_census","text":"","code":"# Match up census tracts for the 1990 and 2000 censuses in Rhode Island if (FALSE) { # \\dontrun{ correspondence <- get_tongfen_correspondence_us_census(datasets = c(\"dec1990\",\"dec2000\"), regions = list(state=\"RI\")) } # }"},{"path":"https://mountainmath.github.io/tongfen/reference/get_tongfen_us_census.html","id":null,"dir":"Reference","previous_headings":"","what":"Get US census data for several censuses on a common geography — get_tongfen_us_census","title":"Get US census data for several censuses on a common geography — get_tongfen_us_census","text":"wraps data acquisition via tidycensus package tongfen common geography single convenience function. Data available 2000, 2010 2020 censuses, Census Bureau retired 1990 API endpoint. tongfen 1990 data, obtain elsewhere combine correspondence table get_tongfen_correspondence_us_census via tongfen_aggregate.","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/get_tongfen_us_census.html","id":"ref-usage","dir":"Reference","previous_headings":"","what":"Usage","title":"Get US census data for several censuses on a common geography — get_tongfen_us_census","text":"","code":"get_tongfen_us_census( regions, meta, level = \"tract\", survey = \"census\", base_geo = NULL, min_area_share = 0.01, sumfile = NULL )"},{"path":"https://mountainmath.github.io/tongfen/reference/get_tongfen_us_census.html","id":"arguments","dir":"Reference","previous_headings":"","what":"Arguments","title":"Get US census data for several censuses on a common geography — get_tongfen_us_census","text":"regions list regions query data . stage, valid list vector states, .e. `regions = list(state=c(\"CA\",\"\"))“ meta metadata variables retrieve level aggregation level return data . stage, valid levels 'tract' 'county subdivision'. survey survey get data , supported options \"census\" base_geo dataset use base geography, example `\"dec2010\"`, one datasets `meta`. Default `NULL`, uses first dataset `meta`. min_area_share minimum share area two geographies common count related, default `0.01`, see get_tongfen_correspondence_us_census. sumfile summary file read variables , either single value used censuses vector named dataset, example `c(dec2010=\"sf1\", dec2020=\"dhc\")`. Default `NULL`, leaves choice tidycensus. Note tidycensus defaults 2020 census PL 94-171 redistricting file, 2020 variables need `sumfile=\"dhc\"`.","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/get_tongfen_us_census.html","id":"value","dir":"Reference","previous_headings":"","what":"Value","title":"Get US census data for several censuses on a common geography — get_tongfen_us_census","text":"sf object (wide form) census variables census year suffix (separated underscore \"_\").","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/get_tongfen_us_census.html","id":"ref-examples","dir":"Reference","previous_headings":"","what":"Examples","title":"Get US census data for several censuses on a common geography — get_tongfen_us_census","text":"","code":"# Get US census data on population and households for 2000 and 2010 censuses on a uniform geography # based on census tracts. if (FALSE) { # \\dontrun{ variables=c(population=\"H011001\",households=\"H013001\") meta <- c(2000,2010) %>% lapply(function(year){ v <- variables %>% setNames(paste0(names(.),\"_\",year)) meta_for_additive_variables(paste0(\"dec\",year),v) }) %>% bind_rows() census_data <- get_tongfen_us_census(regions = list(state=\"CA\"), meta=meta, level=\"tract\") %>% mutate(change=population_2010/households_2010-population_2000/households_2000) } # }"},{"path":"https://mountainmath.github.io/tongfen/reference/meta_for_additive_variables.html","id":null,"dir":"Reference","previous_headings":"","what":"Generate tongfen metadata for additive variables — meta_for_additive_variables","title":"Generate tongfen metadata for additive variables — meta_for_additive_variables","text":"Generates metadata used tongfen_aggregate. Variables need additive like counts.","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/meta_for_additive_variables.html","id":"ref-usage","dir":"Reference","previous_headings":"","what":"Usage","title":"Generate tongfen metadata for additive variables — meta_for_additive_variables","text":"","code":"meta_for_additive_variables(dataset, variables)"},{"path":"https://mountainmath.github.io/tongfen/reference/meta_for_additive_variables.html","id":"arguments","dir":"Reference","previous_headings":"","what":"Arguments","title":"Generate tongfen metadata for additive variables — meta_for_additive_variables","text":"dataset identifier dataset containing variable variables (named) vector additive variables","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/meta_for_additive_variables.html","id":"value","dir":"Reference","previous_headings":"","what":"Value","title":"Generate tongfen metadata for additive variables — meta_for_additive_variables","text":"tibble used tongfen_aggregate","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/meta_for_additive_variables.html","id":"ref-examples","dir":"Reference","previous_headings":"","what":"Examples","title":"Generate tongfen metadata for additive variables — meta_for_additive_variables","text":"","code":"# Get metadata for additive variable Population for the CA16 and CA06 datasets if (FALSE) { # \\dontrun{ meta <- meta_for_additive_variables(c(\"CA06\",\"CA16\"),\"Population\") } # }"},{"path":"https://mountainmath.github.io/tongfen/reference/meta_for_ca_census_vectors.html","id":null,"dir":"Reference","previous_headings":"","what":"Generate metadata from Canadian census vectors — meta_for_ca_census_vectors","title":"Generate metadata from Canadian census vectors — meta_for_ca_census_vectors","text":"Build tibble information aggregate variables given vectors Queries list_census_variables obtain needed information add vectors needed aggregation","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/meta_for_ca_census_vectors.html","id":"ref-usage","dir":"Reference","previous_headings":"","what":"Usage","title":"Generate metadata from Canadian census vectors — meta_for_ca_census_vectors","text":"","code":"meta_for_ca_census_vectors(vectors)"},{"path":"https://mountainmath.github.io/tongfen/reference/meta_for_ca_census_vectors.html","id":"arguments","dir":"Reference","previous_headings":"","what":"Arguments","title":"Generate metadata from Canadian census vectors — meta_for_ca_census_vectors","text":"vectors list variables query","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/meta_for_ca_census_vectors.html","id":"value","dir":"Reference","previous_headings":"","what":"Value","title":"Generate metadata from Canadian census vectors — meta_for_ca_census_vectors","text":"tidy dataframe metadata information requested variables additional variables needed tongfen operations","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/meta_for_ca_census_vectors.html","id":"ref-examples","dir":"Reference","previous_headings":"","what":"Examples","title":"Generate metadata from Canadian census vectors — meta_for_ca_census_vectors","text":"","code":"# Build metadata for vectors if (FALSE) { # \\dontrun{ meta <- meta_for_ca_census_vectors(c(\"v_CA16_4836\",\"v_CA16_4838\",\"v_CA16_4899\")) } # }"},{"path":"https://mountainmath.github.io/tongfen/reference/proportional_reaggregate.html","id":null,"dir":"Reference","previous_headings":"","what":"Dasymetric downsampling — proportional_reaggregate","title":"Dasymetric downsampling — proportional_reaggregate","text":"Proportionally re-aggregate hierarchical data lower-level w.r.t. values *base* variable Also handles cases lower level data may available blinded times filling data higher level Data lower aggregation levels may add accurate aggregate counts. function distributes aggregate level counts proportionally (population) containing lower level geographic regions.","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/proportional_reaggregate.html","id":"ref-usage","dir":"Reference","previous_headings":"","what":"Usage","title":"Dasymetric downsampling — proportional_reaggregate","text":"","code":"proportional_reaggregate( data, parent_data, geo_match, categories, base = \"Population\" )"},{"path":"https://mountainmath.github.io/tongfen/reference/proportional_reaggregate.html","id":"arguments","dir":"Reference","previous_headings":"","what":"Arguments","title":"Dasymetric downsampling — proportional_reaggregate","text":"data base geographic data parent_data Higher level geographic data geo_match named string informing column names match data parent_data categories Vector column names re-aggregate base Column name use proportional weighting re-aggregating, named vector column name category. Categories re-aggregated means set NA reaggregated base data NA values.","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/proportional_reaggregate.html","id":"value","dir":"Reference","previous_headings":"","what":"Value","title":"Dasymetric downsampling — proportional_reaggregate","text":"dataframe downsampled variables parent_data","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/proportional_reaggregate.html","id":"ref-examples","dir":"Reference","previous_headings":"","what":"Examples","title":"Dasymetric downsampling — proportional_reaggregate","text":"","code":"# Proportionally reaggregate visible minority data from dissemination area 2016 # census data to dissemination block geography, proportionally based on dissemination # block population if (FALSE) { # \\dontrun{ regions <- list(CSD=\"5915022\") variables <- cancensus::child_census_vectors(\"v_CA16_3954\") da_data <- cancensus::get_census(\"CA16\",regions=regions, vectors=setNames(variables$vector,variables$label), level=\"DA\") geo_data <- cancensus::get_census(\"CA16\",regions=regions,geo_format=\"sf\",level=\"DB\") db_data <- geo_data %>% proportional_reaggregate(da_data,c(\"DA_UID\"=\"GeoUID\"),variables$label) } # }"},{"path":"https://mountainmath.github.io/tongfen/reference/tongfen_aggregate.html","id":null,"dir":"Reference","previous_headings":"","what":"Perform tongfen according to correspondence — tongfen_aggregate","title":"Perform tongfen according to correspondence — tongfen_aggregate","text":"Aggregate variables specified meta several datasets according correspondence.","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/tongfen_aggregate.html","id":"ref-usage","dir":"Reference","previous_headings":"","what":"Usage","title":"Perform tongfen according to correspondence — tongfen_aggregate","text":"","code":"tongfen_aggregate( data, correspondence, meta = NULL, base_geo = NULL, na.rm = TRUE )"},{"path":"https://mountainmath.github.io/tongfen/reference/tongfen_aggregate.html","id":"arguments","dir":"Reference","previous_headings":"","what":"Arguments","title":"Perform tongfen according to correspondence — tongfen_aggregate","text":"data named list datasets aggregated. names identify datasets, matched `geo_dataset` column `meta` pick aggregation rules labels dataset. Without names, names found `meta`, rules datasets applied variables keep original names correspondence correspondence data gluing datasets meta metadata containing aggregation rules example returned `meta_for_ca_census_vectors` base_geo identifier data element base final geography , uses first data element `NULL` (default), expects `base_geo` element `names(data)`. na.rm logical, determines NA values treated aggregating variables, default `TRUE`","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/tongfen_aggregate.html","id":"value","dir":"Reference","previous_headings":"","what":"Value","title":"Perform tongfen according to correspondence — tongfen_aggregate","text":"aggregated dataset class sf base_geo NULL data type sf tibble otherwise.","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/tongfen_aggregate.html","id":"ref-examples","dir":"Reference","previous_headings":"","what":"Examples","title":"Perform tongfen according to correspondence — tongfen_aggregate","text":"","code":"# aggregate census tract level 2006 and 2016 population data on common geography built through # correspondence from 2006 and 2016 census tracts in the City of Vancouver. if (FALSE) { # \\dontrun{ regions <- list(CSD=\"5915022\") geo1 <- cancensus::get_census(\"CA06\",regions=regions,geo_format='sf',level='CT') geo2 <- cancensus::get_census(\"CA16\",regions=regions,geo_format='sf',level='CT') meta <- meta_for_additive_variables(c(\"CA06\",\"CA16\"),\"Population\") correspondence <- get_tongfen_correspondence_ca_census(geo_datasets=c('CA06','CA16'), regions=regions,level='CT') result <- tongfen_aggregate(list(CA06=geo1 %>% rename(GeoUIDCA06=GeoUID), CA16=geo2 %>% rename(GeoUIDCA16=GeoUID)),correspondence,meta) } # }"},{"path":"https://mountainmath.github.io/tongfen/reference/tongfen_anomaly_joins.html","id":null,"dir":"Reference","previous_headings":"","what":"Determine regions to join to correct for likely geocoding anomalies — tongfen_anomaly_joins","title":"Determine regions to join to correct for likely geocoding anomalies — tongfen_anomaly_joins","text":"Looks regions surprising drops timeline count variable complemented neighbouring region, explained `tongfen_detect_anomalies`, joins . gets repeated joined regions regions left qualify get joined. round region gets joined one region, surprising regions go first. Joining regions trades geographic detail timelines consistent time. parameters control aggressively regions get joined best calibrated data hand, erring side joining regions risks keeping geocoding problems, erring side risks removing real changes needlessly coarsens geography. result can used join regions via `tongfen_join_regions`, update correspondence via `tongfen_join_correspondence`.","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/tongfen_anomaly_joins.html","id":"ref-usage","dir":"Reference","previous_headings":"","what":"Usage","title":"Determine regions to join to correct for likely geocoding anomalies — tongfen_anomaly_joins","text":"","code":"tongfen_anomaly_joins( data, variables, id = \"TongfenID\", neighbours = NULL, rel_scale = 0.25, abs_scale = 200, p = 4, surprise_cutoff = 0.15, total_surprise_cutoff = 0.75, cutoff_fact = 0.6, surprise_reduction_const = 0.15, sum_fact = 0.7 )"},{"path":"https://mountainmath.github.io/tongfen/reference/tongfen_anomaly_joins.html","id":"arguments","dir":"Reference","previous_headings":"","what":"Arguments","title":"Determine regions to join to correct for likely geocoding anomalies — tongfen_anomaly_joins","text":"data data common geography, one row per region, example returned `tongfen_aggregate` `get_tongfen_ca_census`. Needs class sf unless `neighbours` specified. variables names columns holding timeline count variable like population dwellings, temporal order. Changes missing value surprising make surprising changes neighbouring regions, joined regions missing value one regions made . Replace missing values zero beforehand stand regions nothing got counted id name column uniquely identifies regions, default \"TongfenID\" neighbours optional, neighbouring regions table identifiers pairs neighbouring regions first two columns, neighbours list like ones returned `spdep::poly2nb`. default regions intersecting geometries neighbours, can miss neighbours geometries simplified share boundaries . rel_scale relative decrease half way full surprise, default `0.25` 25% drop abs_scale absolute decrease half way full surprise, default `200` p exponent norm used combine surprises across timeline total surprise. Large values focus surprising change, 1 adds surprises changes, default `4` surprise_cutoff changes larger surprise count surprising, default `0.15`. regions least one surprising change candidates total_surprise_cutoff regions larger total surprise candidates, default `0.75` cutoff_fact join regions total surprise joining lower share total surprise candidate region, default `0.6` surprise_reduction_const join regions joining lowers total surprise , default `0.15` sum_fact join regions total surprise joining lower `cutoff_fact * sum_fact` times sum total surprises regions, default `0.7`","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/tongfen_anomaly_joins.html","id":"value","dir":"Reference","previous_headings":"","what":"Value","title":"Determine regions to join to correct for likely geocoding anomalies — tongfen_anomaly_joins","text":"tibble one row region gets joined regions, identifier region, identifier joined region becomes part column named like identifier suffix `_joined`, default `TongfenID_joined`, `round` region first got joined another region. identifier joined region smallest identifier regions made .","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/tongfen_anomaly_joins.html","id":"ref-examples","dir":"Reference","previous_headings":"","what":"Examples","title":"Determine regions to join to correct for likely geocoding anomalies — tongfen_anomaly_joins","text":"","code":"# Correct 2001 through 2021 dissemination area level population timelines in the # City of Vancouver for likely geocoding problems if (FALSE) { # \\dontrun{ datasets <- c(\"CA01\",\"CA06\",\"CA11\",\"CA16\",\"CA21\") meta <- meta_for_additive_variables(datasets,\"Population\") data <- get_tongfen_ca_census(regions=list(CSD=\"5915022\"),meta=meta,level=\"DA\",base_geo=\"CA21\") joins <- tongfen_anomaly_joins(data,paste0(\"Population_\",datasets)) corrected_data <- tongfen_join_regions(data,joins,meta) } # }"},{"path":"https://mountainmath.github.io/tongfen/reference/tongfen_ca_census_ct.html","id":null,"dir":"Reference","previous_headings":"","what":"Canadian census CT level tongfen via identifier matching — tongfen_ca_census_ct","title":"Canadian census CT level tongfen via identifier matching — tongfen_ca_census_ct","text":"Aggregate variables common CTs, returns data2 new tiling matching data1 geography","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/tongfen_ca_census_ct.html","id":"ref-usage","dir":"Reference","previous_headings":"","what":"Usage","title":"Canadian census CT level tongfen via identifier matching — tongfen_ca_census_ct","text":"","code":"tongfen_ca_census_ct( data1, data2, data2_sum_vars, data2_group_vars = c(), na.rm = TRUE )"},{"path":"https://mountainmath.github.io/tongfen/reference/tongfen_ca_census_ct.html","id":"arguments","dir":"Reference","previous_headings":"","what":"Arguments","title":"Canadian census CT level tongfen via identifier matching — tongfen_ca_census_ct","text":"data1 cancensus CT level dataset year1 < year2 serve base common geography data2 cancensus CT level dataset year2 aggregated common geography data2_sum_vars vector variable names summed aggregating geographies data2_group_vars optional vector grouping variables na.rm optional parameter remove NA values summing, default = `TRUE`","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/tongfen_ca_census_ct.html","id":"value","dir":"Reference","previous_headings":"","what":"Value","title":"Canadian census CT level tongfen via identifier matching — tongfen_ca_census_ct","text":"`data2` variables `data2_sum_vars` aggregated common geography matching `data1`, identified `GeoUID` `data1`","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/tongfen_detect_anomalies.html","id":null,"dir":"Reference","previous_headings":"","what":"Detect likely geocoding anomalies in timelines on a common geography — tongfen_detect_anomalies","title":"Detect likely geocoding anomalies in timelines on a common geography — tongfen_detect_anomalies","text":"TongFen good geocoding assigned underlying data geographic regions first place. Geocoding varies time, dwelling units, people living , can get assigned different neighbouring regions different years. timeline common geography shows surprising drop one region offset corresponding jump neighbouring region. function lists candidate regions surprising drops given count variable, together neighbouring region takes away surprise joined. Use check possible problems calibrate parameters joining regions `tongfen_anomaly_joins`. surprising drops due geocoding problems, drop complemented neighbouring region likely real. surprise change two consecutive years ranges 0 1. decreases surprising, surprise product surprise relative absolute decrease, takes decrease large relative absolute terms surprising. total surprise region `p`-norm surprises across changes timeline. candidate region neighbour flagged joining joining reduces total surprise candidate region `cutoff_fact` times total surprise, reduces `surprise_reduction_const`, reduces `cutoff_fact * sum_fact` times sum total surprises regions. keep comparable surprise joining computed change joined regions relative counts candidate region alone.","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/tongfen_detect_anomalies.html","id":"ref-usage","dir":"Reference","previous_headings":"","what":"Usage","title":"Detect likely geocoding anomalies in timelines on a common geography — tongfen_detect_anomalies","text":"","code":"tongfen_detect_anomalies( data, variables, id = \"TongfenID\", neighbours = NULL, rel_scale = 0.25, abs_scale = 200, p = 4, surprise_cutoff = 0.15, total_surprise_cutoff = 0.75, cutoff_fact = 0.6, surprise_reduction_const = 0.15, sum_fact = 0.7 )"},{"path":"https://mountainmath.github.io/tongfen/reference/tongfen_detect_anomalies.html","id":"arguments","dir":"Reference","previous_headings":"","what":"Arguments","title":"Detect likely geocoding anomalies in timelines on a common geography — tongfen_detect_anomalies","text":"data data common geography, one row per region, example returned `tongfen_aggregate` `get_tongfen_ca_census`. Needs class sf unless `neighbours` specified. variables names columns holding timeline count variable like population dwellings, temporal order. Changes missing value surprising make surprising changes neighbouring regions, joined regions missing value one regions made . Replace missing values zero beforehand stand regions nothing got counted id name column uniquely identifies regions, default \"TongfenID\" neighbours optional, neighbouring regions table identifiers pairs neighbouring regions first two columns, neighbours list like ones returned `spdep::poly2nb`. default regions intersecting geometries neighbours, can miss neighbours geometries simplified share boundaries . rel_scale relative decrease half way full surprise, default `0.25` 25% drop abs_scale absolute decrease half way full surprise, default `200` p exponent norm used combine surprises across timeline total surprise. Large values focus surprising change, 1 adds surprises changes, default `4` surprise_cutoff changes larger surprise count surprising, default `0.15`. regions least one surprising change candidates total_surprise_cutoff regions larger total surprise candidates, default `0.75` cutoff_fact join regions total surprise joining lower share total surprise candidate region, default `0.6` surprise_reduction_const join regions joining lowers total surprise , default `0.15` sum_fact join regions total surprise joining lower `cutoff_fact * sum_fact` times sum total surprises regions, default `0.7`","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/tongfen_detect_anomalies.html","id":"value","dir":"Reference","previous_headings":"","what":"Value","title":"Detect likely geocoding anomalies in timelines on a common geography — tongfen_detect_anomalies","text":"tibble one row candidate region, surprising first, identifier region, number surprising changes `surprise_count`, total surprise `surprise_total`, `period` surprising change, identifier `neighbour` takes away surprise, total surprise `surprise_total_joined` joining `join` indicating regions qualify get joined.","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/tongfen_detect_anomalies.html","id":"ref-examples","dir":"Reference","previous_headings":"","what":"Examples","title":"Detect likely geocoding anomalies in timelines on a common geography — tongfen_detect_anomalies","text":"","code":"# Check 2001 through 2021 dissemination area level population timelines in the # City of Vancouver for possible geocoding problems if (FALSE) { # \\dontrun{ datasets <- c(\"CA01\",\"CA06\",\"CA11\",\"CA16\",\"CA21\") meta <- meta_for_additive_variables(datasets,\"Population\") data <- get_tongfen_ca_census(regions=list(CSD=\"5915022\"),meta=meta,level=\"DA\",base_geo=\"CA21\") anomalies <- tongfen_detect_anomalies(data,paste0(\"Population_\",datasets)) } # }"},{"path":"https://mountainmath.github.io/tongfen/reference/tongfen_estimate.html","id":null,"dir":"Reference","previous_headings":"","what":"Estimate variable values for custom geography — tongfen_estimate","title":"Estimate variable values for custom geography — tongfen_estimate","text":"Estimates data source geometry onto target geometry using area-weighted interpolation. metadata specifies data aggregated, \"additive\" data like population counts summed proportionally area intersection, \"averages\" need additive \"parent\" count variables estimate weighted averages.","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/tongfen_estimate.html","id":"ref-usage","dir":"Reference","previous_headings":"","what":"Usage","title":"Estimate variable values for custom geography — tongfen_estimate","text":"","code":"tongfen_estimate(target, source, meta, na.rm = FALSE)"},{"path":"https://mountainmath.github.io/tongfen/reference/tongfen_estimate.html","id":"arguments","dir":"Reference","previous_headings":"","what":"Arguments","title":"Estimate variable values for custom geography — tongfen_estimate","text":"target custom geography estimate values source input geography values meta metadata variable aggregation, see `meta_for_additive_variables` `meta_for_ca_census_vectors` information construct metadata. na.rm remove NA values aggregating, default FALSE","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/tongfen_estimate.html","id":"value","dir":"Reference","previous_headings":"","what":"Value","title":"Estimate variable values for custom geography — tongfen_estimate","text":"`target` estimated quantities `source` specified `meta`, regions `target` overlap `source` `NA` values. Columns `target` name variables estimated.","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/tongfen_estimate.html","id":"ref-examples","dir":"Reference","previous_headings":"","what":"Examples","title":"Estimate variable values for custom geography — tongfen_estimate","text":"","code":"# Estimate 2006 Population in the City of Vancouver dissemination ares on 2016 census geographies if (FALSE) { # \\dontrun{ geo1 <- cancensus::get_census(\"CA06\",regions=list(CSD=\"5915022\"),geo_format='sf',level='DA') geo2 <- cancensus::get_census(\"CA16\",regions=list(CSD=\"5915022\"),geo_format='sf',level='DA') meta <- meta_for_additive_variables(\"CA06\",\"Population\") result <- tongfen_estimate(geo2 %>% rename(Population_2016=Population),geo1,meta) } # }"},{"path":"https://mountainmath.github.io/tongfen/reference/tongfen_estimate_ca_census.html","id":null,"dir":"Reference","previous_headings":"","what":"Tongfen estimate data for given geometry — tongfen_estimate_ca_census","title":"Tongfen estimate data for given geometry — tongfen_estimate_ca_census","text":"Estimates values given census vectors given geometry using data specified level range. wrapper around `cancensus::get_intersecting_geometries` `tongfen_estimate`, optionally downsampling via `proportional_reaggregate`, streamline estimating Canadian census data custom geographies.","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/tongfen_estimate_ca_census.html","id":"ref-usage","dir":"Reference","previous_headings":"","what":"Usage","title":"Tongfen estimate data for given geometry — tongfen_estimate_ca_census","text":"","code":"tongfen_estimate_ca_census( geometry, meta, level, intersection_level = level, downsample_level = NULL, na.rm = FALSE, quiet = FALSE )"},{"path":"https://mountainmath.github.io/tongfen/reference/tongfen_estimate_ca_census.html","id":"arguments","dir":"Reference","previous_headings":"","what":"Arguments","title":"Tongfen estimate data for given geometry — tongfen_estimate_ca_census","text":"geometry geometry meta metadata census variables aggregate, example returned `meta_for_ca_census_vectors`. point function accepts variables census geography year. expand also allow estimates across multiple census geography years, requires attention detail. recommended apply due caution running function separately across several census geography years purpose comparing data across time naive application can lead systematic biases. level level use tongfen intersection_level level use geometry intersection, different tongfen level meta_for_ca_census_vectors. can set higher aggregation level conserve API points `get_intersecting_geometries` call. downsample_level default `NULL`, can geographic level lower `level`, case data downsamples geography level proportionally using value `downsample` column (must supplied) `meta` argument intersecting geometries. can lead accurate results. point allowed variables `downsample` column `meta` \"Population\", \"Households\" \"Dwellings\", can one variables. na.rm deal NA values, default FALSE. quiet suppress progress messages","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/tongfen_estimate_ca_census.html","id":"value","dir":"Reference","previous_headings":"","what":"Value","title":"Tongfen estimate data for given geometry — tongfen_estimate_ca_census","text":"`geometry` estimated values census variables specified `meta`","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/tongfen_estimate_ca_census.html","id":"ref-examples","dir":"Reference","previous_headings":"","what":"Examples","title":"Tongfen estimate data for given geometry — tongfen_estimate_ca_census","text":"","code":"# Estimate the 2016 population within 1 km of Toronto City Hall from dissemination area level # census data if (FALSE) { # \\dontrun{ toronto_city_hall <- sf::st_point(c(-79.3839,43.6534)) %>% sf::st_sfc(crs=4326) %>% sf::st_transform(3348) %>% sf::st_buffer(1000) %>% sf::st_sf() meta <- meta_for_additive_variables(\"CA16\",\"Population\") data <- tongfen_estimate_ca_census(toronto_city_hall,meta,level=\"DA\",intersection_level=\"CT\") print(paste0(\"Approximately \",scales::comma(data$Population,accuracy=100), \" people live within a 1 km radius of Toronto City Hall.\")) } # }"},{"path":"https://mountainmath.github.io/tongfen/reference/tongfen_join_correspondence.html","id":null,"dir":"Reference","previous_headings":"","what":"Join regions in a correspondence — tongfen_join_correspondence","title":"Join regions in a correspondence — tongfen_join_correspondence","text":"Updates correspondence given regions joined, example correct likely geocoding anomalies determined `tongfen_anomaly_joins`. updated correspondence can used `tongfen_aggregate` aggregate data coarser common geography, works variables `tongfen_aggregate` can deal data part detecting anomalies.","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/tongfen_join_correspondence.html","id":"ref-usage","dir":"Reference","previous_headings":"","what":"Usage","title":"Join regions in a correspondence — tongfen_join_correspondence","text":"","code":"tongfen_join_correspondence(correspondence, joins)"},{"path":"https://mountainmath.github.io/tongfen/reference/tongfen_join_correspondence.html","id":"arguments","dir":"Reference","previous_headings":"","what":"Arguments","title":"Join regions in a correspondence — tongfen_join_correspondence","text":"correspondence correspondence table columns unique geographic identifiers geographies TongfenID TongfenUID, example returned `estimate_tongfen_correspondence` `get_tongfen_correspondence_ca_census` joins table regions join returned `tongfen_anomaly_joins`, columns `TongfenID` `TongfenID_joined`","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/tongfen_join_correspondence.html","id":"value","dir":"Reference","previous_headings":"","what":"Value","title":"Join regions in a correspondence — tongfen_join_correspondence","text":"correspondence updated TongfenID TongfenUID regions got joined. correspondence TongfenMethod column \"anomaly\" gets added method regions got joined.","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/tongfen_join_correspondence.html","id":"ref-examples","dir":"Reference","previous_headings":"","what":"Examples","title":"Join regions in a correspondence — tongfen_join_correspondence","text":"","code":"# Correct for likely geocoding problems in dissemination area level population timelines # and use the updated correspondence to aggregate data on the corrected common geography if (FALSE) { # \\dontrun{ regions <- list(CSD=\"5915022\") datasets <- c(\"CA01\",\"CA06\",\"CA11\",\"CA16\",\"CA21\") meta <- meta_for_additive_variables(datasets,\"Population\") data <- get_tongfen_ca_census(regions=regions,meta=meta,level=\"DA\",base_geo=\"CA21\") joins <- tongfen_anomaly_joins(data,paste0(\"Population_\",datasets)) correspondence <- get_tongfen_correspondence_ca_census(geo_datasets=datasets, regions=regions,level=\"DA\") %>% tongfen_join_correspondence(joins) } # }"},{"path":"https://mountainmath.github.io/tongfen/reference/tongfen_join_regions.html","id":null,"dir":"Reference","previous_headings":"","what":"Join regions in data on a common geography — tongfen_join_regions","title":"Join regions in data on a common geography — tongfen_join_regions","text":"Joins regions data already aggregated common geography, example correct likely geocoding anomalies determined `tongfen_anomaly_joins`. data, geometries data class sf, regions get joined aggregated, regions left . Variables aggregated according metadata, numeric variables part metadata assumed additive. Variables additive, like averages, can aggregated parent variable part data. case use `tongfen_join_correspondence` update correspondence data built aggregate original data `tongfen_aggregate`.","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/tongfen_join_regions.html","id":"ref-usage","dir":"Reference","previous_headings":"","what":"Usage","title":"Join regions in data on a common geography — tongfen_join_regions","text":"","code":"tongfen_join_regions(data, joins, meta = NULL, id = \"TongfenID\", na.rm = TRUE)"},{"path":"https://mountainmath.github.io/tongfen/reference/tongfen_join_regions.html","id":"arguments","dir":"Reference","previous_headings":"","what":"Arguments","title":"Join regions in data on a common geography — tongfen_join_regions","text":"data data common geography, one row per region, example returned `tongfen_aggregate` `get_tongfen_ca_census` joins table regions join returned `tongfen_anomaly_joins`, identifier region identifier joined region becomes part column named like identifier suffix `_joined` meta optional metadata containing aggregation rules example returned `meta_for_ca_census_vectors`, variables matched label. Numeric variables part metadata treated additive, `NULL` (default) case numeric variables id name column uniquely identifies regions, default \"TongfenID\" na.rm logical, determines NA values treated aggregating variables, default `TRUE`","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/tongfen_join_regions.html","id":"value","dir":"Reference","previous_headings":"","what":"Value","title":"Join regions in data on a common geography — tongfen_join_regions","text":"data regions joined. Joined regions take place identifier region smallest identifier among regions made . Variables numeric part metadata `NA` joined regions.","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/tongfen_join_regions.html","id":"ref-examples","dir":"Reference","previous_headings":"","what":"Examples","title":"Join regions in data on a common geography — tongfen_join_regions","text":"","code":"# Correct 2001 through 2021 dissemination area level population timelines in the # City of Vancouver for likely geocoding problems if (FALSE) { # \\dontrun{ datasets <- c(\"CA01\",\"CA06\",\"CA11\",\"CA16\",\"CA21\") meta <- meta_for_additive_variables(datasets,\"Population\") data <- get_tongfen_ca_census(regions=list(CSD=\"5915022\"),meta=meta,level=\"DA\",base_geo=\"CA21\") joins <- tongfen_anomaly_joins(data,paste0(\"Population_\",datasets)) corrected_data <- tongfen_join_regions(data,joins,meta) } # }"},{"path":"https://mountainmath.github.io/tongfen/reference/tongfen_tag_largest_overlap.html","id":null,"dir":"Reference","previous_headings":"","what":"Tag regions by largest overlap — tongfen_tag_largest_overlap","title":"Tag regions by largest overlap — tongfen_tag_largest_overlap","text":"tags regions `source` `target_id` region `target` largest overlap","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/tongfen_tag_largest_overlap.html","id":"ref-usage","dir":"Reference","previous_headings":"","what":"Usage","title":"Tag regions by largest overlap — tongfen_tag_largest_overlap","text":"","code":"tongfen_tag_largest_overlap(source, target, target_id)"},{"path":"https://mountainmath.github.io/tongfen/reference/tongfen_tag_largest_overlap.html","id":"arguments","dir":"Reference","previous_headings":"","what":"Arguments","title":"Tag regions by largest overlap — tongfen_tag_largest_overlap","text":"source input geography target custom geography target_id name column `target` table unique id (character)","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/tongfen_tag_largest_overlap.html","id":"value","dir":"Reference","previous_headings":"","what":"Value","title":"Tag regions by largest overlap — tongfen_tag_largest_overlap","text":"`source` extra column name given `target_id` column `...overlap_fraction` proportion area source region overlaps region `target` id","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/tongfen_tag_largest_overlap.html","id":"ref-examples","dir":"Reference","previous_headings":"","what":"Examples","title":"Tag regions by largest overlap — tongfen_tag_largest_overlap","text":"","code":"# Tag 2016 dissemination areas in the City of Vancouver by the 2006 census tract they overlap # the most with if (FALSE) { # \\dontrun{ geo1 <- cancensus::get_census(\"CA06\",regions=list(CSD=\"5915022\"),geo_format='sf',level='CT') geo2 <- cancensus::get_census(\"CA16\",regions=list(CSD=\"5915022\"),geo_format='sf',level='DA') result <- tongfen_tag_largest_overlap(geo2,geo1 %>% select(CT_2006=GeoUID),\"CT_2006\") } # }"},{"path":"https://mountainmath.github.io/tongfen/reference/vancouver_elections_data_2015.html","id":null,"dir":"Reference","previous_headings":"","what":"A dataset with polling station votes data from the 2015 federal election in the Vancouver area — vancouver_elections_data_2015","title":"A dataset with polling station votes data from the 2015 federal election in the Vancouver area — vancouver_elections_data_2015","text":"dataset polling station votes data 2015 federal election Vancouver area","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/vancouver_elections_data_2015.html","id":"references","dir":"Reference","previous_headings":"","what":"References","title":"A dataset with polling station votes data from the 2015 federal election in the Vancouver area — vancouver_elections_data_2015","text":"https://www.elections.ca/content.aspx?section=res&dir=rep/&document=index&lang=e#42GE","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/vancouver_elections_data_2015.html","id":"author","dir":"Reference","previous_headings":"","what":"Author","title":"A dataset with polling station votes data from the 2015 federal election in the Vancouver area — vancouver_elections_data_2015","text":"Elections Canada","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/vancouver_elections_data_2019.html","id":null,"dir":"Reference","previous_headings":"","what":"A dataset with polling station votes data from the 2019 federal election in the Vancouver area — vancouver_elections_data_2019","title":"A dataset with polling station votes data from the 2019 federal election in the Vancouver area — vancouver_elections_data_2019","text":"dataset polling station votes data 2019 federal election Vancouver area","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/vancouver_elections_data_2019.html","id":"references","dir":"Reference","previous_headings":"","what":"References","title":"A dataset with polling station votes data from the 2019 federal election in the Vancouver area — vancouver_elections_data_2019","text":"https://www.elections.ca/content.aspx?section=res&dir=rep/&document=index&lang=e#43GE","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/vancouver_elections_data_2019.html","id":"author","dir":"Reference","previous_headings":"","what":"Author","title":"A dataset with polling station votes data from the 2019 federal election in the Vancouver area — vancouver_elections_data_2019","text":"Elections Canada","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/vancouver_elections_geos_2015.html","id":null,"dir":"Reference","previous_headings":"","what":"A dataset with polling district geographies from the 2015 federal election in the Vancouver area — vancouver_elections_geos_2015","title":"A dataset with polling district geographies from the 2015 federal election in the Vancouver area — vancouver_elections_geos_2015","text":"dataset polling district geographies 2015 federal election Vancouver area","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/vancouver_elections_geos_2015.html","id":"references","dir":"Reference","previous_headings":"","what":"References","title":"A dataset with polling district geographies from the 2015 federal election in the Vancouver area — vancouver_elections_geos_2015","text":"https://www.elections.ca/content.aspx?section=res&dir=rep/&document=index&lang=e#42GE","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/vancouver_elections_geos_2015.html","id":"author","dir":"Reference","previous_headings":"","what":"Author","title":"A dataset with polling district geographies from the 2015 federal election in the Vancouver area — vancouver_elections_geos_2015","text":"Elections Canada","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/vancouver_elections_geos_2019.html","id":null,"dir":"Reference","previous_headings":"","what":"A dataset with polling district geographies from the 2019 federal election in the Vancouver area — vancouver_elections_geos_2019","title":"A dataset with polling district geographies from the 2019 federal election in the Vancouver area — vancouver_elections_geos_2019","text":"dataset polling district geographies 2019 federal election Vancouver area","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/vancouver_elections_geos_2019.html","id":"references","dir":"Reference","previous_headings":"","what":"References","title":"A dataset with polling district geographies from the 2019 federal election in the Vancouver area — vancouver_elections_geos_2019","text":"https://www.elections.ca/content.aspx?section=res&dir=rep/&document=index&lang=e#43GE","code":""},{"path":"https://mountainmath.github.io/tongfen/reference/vancouver_elections_geos_2019.html","id":"author","dir":"Reference","previous_headings":"","what":"Author","title":"A dataset with polling district geographies from the 2019 federal election in the Vancouver area — vancouver_elections_geos_2019","text":"Elections Canada","code":""},{"path":"https://mountainmath.github.io/tongfen/news/index.html","id":"tongfen-032","dir":"Changelog","previous_headings":"","what":"tongfen 0.3.2","title":"tongfen 0.3.2","text":"Fix compatibility issue changes {sf} package reliable GitHub action CRAN checks","code":""},{"path":[]},{"path":"https://mountainmath.github.io/tongfen/news/index.html","id":"major-changes-0-3-2","dir":"Changelog","previous_headings":"","what":"Major changes","title":"tongfen 0.3.2","text":"Added tongfen_estimate_ca_census function new CensusMapper endpoint, tying new {cancensus} functionality. ## Minor changes Custom implementation tongfen_estimate finer control Fix compatibility issue changes {sf} package","code":""},{"path":"https://mountainmath.github.io/tongfen/news/index.html","id":"tongfen-03","dir":"Changelog","previous_headings":"","what":"tongfen 0.3","title":"tongfen 0.3","text":"CRAN release: 2020-11-04","code":""},{"path":"https://mountainmath.github.io/tongfen/news/index.html","id":"major-changes-0-3","dir":"Changelog","previous_headings":"","what":"Major changes","title":"tongfen 0.3","text":"Initial release","code":""}]
diff --git a/docs/sitemap.xml b/docs/sitemap.xml
index 4785765..208aa23 100644
--- a/docs/sitemap.xml
+++ b/docs/sitemap.xml
@@ -6,6 +6,7 @@
https://mountainmath.github.io/tongfen/articles/polling_districts.htmlhttps://mountainmath.github.io/tongfen/articles/tongfen-ca-estimate.htmlhttps://mountainmath.github.io/tongfen/articles/tongfen.html
+https://mountainmath.github.io/tongfen/articles/tongfen_anomalies.htmlhttps://mountainmath.github.io/tongfen/articles/tongfen_ca.htmlhttps://mountainmath.github.io/tongfen/articles/tongfen_us.htmlhttps://mountainmath.github.io/tongfen/authors.html
@@ -31,9 +32,13 @@
https://mountainmath.github.io/tongfen/reference/meta_for_ca_census_vectors.htmlhttps://mountainmath.github.io/tongfen/reference/proportional_reaggregate.htmlhttps://mountainmath.github.io/tongfen/reference/tongfen_aggregate.html
+https://mountainmath.github.io/tongfen/reference/tongfen_anomaly_joins.htmlhttps://mountainmath.github.io/tongfen/reference/tongfen_ca_census_ct.html
+https://mountainmath.github.io/tongfen/reference/tongfen_detect_anomalies.htmlhttps://mountainmath.github.io/tongfen/reference/tongfen_estimate.htmlhttps://mountainmath.github.io/tongfen/reference/tongfen_estimate_ca_census.html
+https://mountainmath.github.io/tongfen/reference/tongfen_join_correspondence.html
+https://mountainmath.github.io/tongfen/reference/tongfen_join_regions.htmlhttps://mountainmath.github.io/tongfen/reference/tongfen_tag_largest_overlap.htmlhttps://mountainmath.github.io/tongfen/reference/vancouver_elections_data_2015.htmlhttps://mountainmath.github.io/tongfen/reference/vancouver_elections_data_2019.html
diff --git a/man/add_census_ca_base_variables.Rd b/man/add_census_ca_base_variables.Rd
index c1a5406..371400d 100644
--- a/man/add_census_ca_base_variables.Rd
+++ b/man/add_census_ca_base_variables.Rd
@@ -2,12 +2,12 @@
% Please edit documentation in R/tongfen_ca.R
\name{add_census_ca_base_variables}
\alias{add_census_ca_base_variables}
-\title{Generate metadata from Candian census vectors}
+\title{Generate metadata from Canadian census vectors}
\usage{
add_census_ca_base_variables(meta)
}
\arguments{
-\item{meta}{ribble with metadata as for example provided by `meta_for_ca_census_vectors`}
+\item{meta}{tibble with metadata as for example provided by `meta_for_ca_census_vectors`}
}
\value{
tibble with metadata
diff --git a/man/check_tongfen_areas.Rd b/man/check_tongfen_areas.Rd
index d03414a..99a35e6 100644
--- a/man/check_tongfen_areas.Rd
+++ b/man/check_tongfen_areas.Rd
@@ -2,12 +2,12 @@
% Please edit documentation in R/tongfen.R
\name{check_tongfen_areas}
\alias{check_tongfen_areas}
-\title{Check geographic integrety}
+\title{Check geographic integrity}
\usage{
check_tongfen_areas(data, correspondence)
}
\arguments{
-\item{data}{alist of geogrpahic data of class sf}
+\item{data}{a list of geographic data of class sf}
\item{correspondence}{Correspondence table with columns the unique geographic identifiers for each of the
geographies and the TongfenID (and optionally TongfenUID and TongfenMethod)
diff --git a/man/check_tongfen_single_areas.Rd b/man/check_tongfen_single_areas.Rd
index f20795f..ec7401d 100644
--- a/man/check_tongfen_single_areas.Rd
+++ b/man/check_tongfen_single_areas.Rd
@@ -2,7 +2,7 @@
% Please edit documentation in R/tonfen_deprecated.R
\name{check_tongfen_single_areas}
\alias{check_tongfen_single_areas}
-\title{Check geographic integrety}
+\title{Check geographic integrity}
\usage{
check_tongfen_single_areas(geo1, geo2, correspondence)
}
diff --git a/man/estimate_tongfen_correspondence.Rd b/man/estimate_tongfen_correspondence.Rd
index 0322925..56576a1 100644
--- a/man/estimate_tongfen_correspondence.Rd
+++ b/man/estimate_tongfen_correspondence.Rd
@@ -2,7 +2,7 @@
% Please edit documentation in R/tongfen.R
\name{estimate_tongfen_correspondence}
\alias{estimate_tongfen_correspondence}
-\title{Generate togfen correspondence for list of geographies}
+\title{Generate tongfen correspondence for list of geographies}
\usage{
estimate_tongfen_correspondence(
data,
diff --git a/man/estimate_tongfen_single_correspondence.Rd b/man/estimate_tongfen_single_correspondence.Rd
index 5a44c3e..a145363 100644
--- a/man/estimate_tongfen_single_correspondence.Rd
+++ b/man/estimate_tongfen_single_correspondence.Rd
@@ -2,7 +2,7 @@
% Please edit documentation in R/tongfen.R
\name{estimate_tongfen_single_correspondence}
\alias{estimate_tongfen_single_correspondence}
-\title{Generate togfen correspondence for two geographies}
+\title{Generate tongfen correspondence for two geographies}
\usage{
estimate_tongfen_single_correspondence(
geo1,
diff --git a/man/get_correspondence_ca_census_for.Rd b/man/get_correspondence_ca_census_for.Rd
index 90bec0d..25a37ec 100644
--- a/man/get_correspondence_ca_census_for.Rd
+++ b/man/get_correspondence_ca_census_for.Rd
@@ -18,5 +18,5 @@ tibble with correspondence table`spanning all years
}
\description{
\lifecycle{deprecated}
-Joins the StatCan correspodence files for several census years
+Joins the StatCan correspondence files for several census years
}
diff --git a/man/get_single_correspondence_ca_census_for.Rd b/man/get_single_correspondence_ca_census_for.Rd
index 3970fba..bd4c0e4 100644
--- a/man/get_single_correspondence_ca_census_for.Rd
+++ b/man/get_single_correspondence_ca_census_for.Rd
@@ -22,4 +22,9 @@ tibble with correspondence table`
}
\description{
\lifecycle{maturing}
+
+The correspondence files are downloaded from a mirror of the Statistics Canada correspondence files
+and cached in the tongfen cache directory. The cached files are checked against the mirror once per session
+and get downloaded again if they changed. The location of the mirror can be changed via the
+`tongfen.statcan_correspondence_url` option.
}
diff --git a/man/get_tongfen_ca_census.Rd b/man/get_tongfen_ca_census.Rd
index b5d388c..4962ec0 100644
--- a/man/get_tongfen_ca_census.Rd
+++ b/man/get_tongfen_ca_census.Rd
@@ -2,7 +2,7 @@
% Please edit documentation in R/tongfen_ca.R
\name{get_tongfen_ca_census}
\alias{get_tongfen_ca_census}
-\title{Togfen data from several Canadian censuses}
+\title{Tongfen data from several Canadian censuses}
\usage{
get_tongfen_ca_census(
regions,
@@ -21,7 +21,7 @@ get_tongfen_ca_census(
\arguments{
\item{regions}{census region list, should be inclusive list of GeoUIDs across censuses}
-\item{meta}{metadata for the census veraiables to aggregate, for example as returned
+\item{meta}{metadata for the census variables to aggregate, for example as returned
by \code{meta_for_ca_census_vectors}.}
\item{level}{aggregation level to return data on (default is "CT")}
@@ -38,7 +38,7 @@ any geographic data}
\item{na.rm}{logical, determines how NA values should be treated when aggregating variables,
default is `FALSE`}
-\item{tolerance}{tolerance for `estimate_tongen_correspondence` in metres, default value is 50 metres,
+\item{tolerance}{tolerance for `estimate_tongfen_correspondence` in metres, default value is 50 metres,
only used when method is 'estimate' or 'identifier'}
\item{quiet}{suppress download progress output, default is `FALSE`}
@@ -56,7 +56,7 @@ dataframe with variables on common geography
\description{
\lifecycle{maturing}
-Get data from several Candian censuses on a common geography. Requires sf and cancensus package to be available
+Get data from several Canadian censuses on a common geography. Requires sf and cancensus package to be available
}
\examples{
# Get rent data for census years 2001 through 2016
diff --git a/man/get_tongfen_correspondence_ca_census.Rd b/man/get_tongfen_correspondence_ca_census.Rd
index ad8d6c8..5d0cb82 100644
--- a/man/get_tongfen_correspondence_ca_census.Rd
+++ b/man/get_tongfen_correspondence_ca_census.Rd
@@ -28,7 +28,7 @@ this method only works for "DB", "DA" and "CT" levels.
* "estimate" uses `estimate_tongfen_correspondence` to build up the common geography from scratch based on geographies.
* "identifier" assumes regions with identical geographic identifier are identical, and builds up the the correspondence for regions with unmatched geographic identifiers.}
-\item{tolerance}{tolerance for `estimate_tongen_correspondence` in metres, default value is 50 metres,
+\item{tolerance}{tolerance for `estimate_tongfen_correspondence` in metres, default value is 50 metres,
only used when method is 'estimate' or 'identifier'}
\item{quiet}{suppress download progress output, default is `FALSE`}
@@ -44,7 +44,7 @@ dataframe with the multi-census correspondence file
\description{
\lifecycle{maturing}
-Get correspondence file for several Candian censuses on a common geography. Requires sf and cancensus package to be available
+Get correspondence file for several Canadian censuses on a common geography. Requires sf and cancensus package to be available
}
\examples{
# Get correspondance files between CTs in 2006 and 2016 censuses in Vancouver CMA
diff --git a/man/get_tongfen_correspondence_us_census.Rd b/man/get_tongfen_correspondence_us_census.Rd
index 71a3170..96f22eb 100644
--- a/man/get_tongfen_correspondence_us_census.Rd
+++ b/man/get_tongfen_correspondence_us_census.Rd
@@ -31,7 +31,8 @@ geographies at the risk of separating regions that did change. No region is ever
if all of its parts are slivers its largest part is kept.}
\item{cache_path}{optional path to cache the relationship files in, defaults to the
-`tongfen.cache_path` option and falls back to a temporary directory}
+`tongfen.cache_path` option. If that is not set the `tongfen.cache_path` environment variable
+and the `custom_data_path` option are used, falling back to a temporary directory}
}
\value{
tibble with one row per census geography, a GEOID column for each requested census,
diff --git a/man/get_tongfen_us_census.Rd b/man/get_tongfen_us_census.Rd
index d6235cd..f093348 100644
--- a/man/get_tongfen_us_census.Rd
+++ b/man/get_tongfen_us_census.Rd
@@ -2,7 +2,7 @@
% Please edit documentation in R/tongfen_us.R
\name{get_tongfen_us_census}
\alias{get_tongfen_us_census}
-\title{Get US census data for 2000 and 2010 census on common census tract based geography}
+\title{Get US census data for several censuses on a common geography}
\usage{
get_tongfen_us_census(
regions,
@@ -24,7 +24,8 @@ valid list is a vector of states, i.e. `regions = list(state=c("CA","OR"))``}
\item{survey}{survey to get data for, supported options is "census"}
-\item{base_geo}{census year to use as base geography, default is `2010`.}
+\item{base_geo}{dataset to use as base geography, for example `"dec2010"`, has to be one of
+the datasets in `meta`. Default is `NULL`, which uses the first dataset in `meta`.}
\item{min_area_share}{minimum share of area two geographies have to have in common to count
as related, default is `0.01`, see \code{\link{get_tongfen_correspondence_us_census}}.}
@@ -35,7 +36,7 @@ is `NULL`, which leaves the choice to tidycensus. Note that tidycensus defaults
census to the PL 94-171 redistricting file, most 2020 variables need `sumfile="dhc"`.}
}
\value{
-sf object with (wide form) census variables with census year as suffix (separated by underdcore "_").
+sf object with (wide form) census variables with census year as suffix (separated by underscore "_").
}
\description{
\lifecycle{maturing}
diff --git a/man/meta_for_additive_variables.Rd b/man/meta_for_additive_variables.Rd
index 468e53f..42ca9af 100644
--- a/man/meta_for_additive_variables.Rd
+++ b/man/meta_for_additive_variables.Rd
@@ -7,9 +7,9 @@
meta_for_additive_variables(dataset, variables)
}
\arguments{
-\item{dataset}{identifier for the dataset contianing the variable}
+\item{dataset}{identifier for the dataset containing the variable}
-\item{variables}{(named) vecotor with additive variables}
+\item{variables}{(named) vector with additive variables}
}
\value{
a tibble to be used in tongfen_aggregate
diff --git a/man/meta_for_ca_census_vectors.Rd b/man/meta_for_ca_census_vectors.Rd
index 3fbaa35..b911536 100644
--- a/man/meta_for_ca_census_vectors.Rd
+++ b/man/meta_for_ca_census_vectors.Rd
@@ -2,7 +2,7 @@
% Please edit documentation in R/tongfen_ca.R
\name{meta_for_ca_census_vectors}
\alias{meta_for_ca_census_vectors}
-\title{Generate metadata from Candian census vectors}
+\title{Generate metadata from Canadian census vectors}
\usage{
meta_for_ca_census_vectors(vectors)
}
@@ -22,6 +22,6 @@ Queries list_census_variables to obtain needed information and add in vectors ne
\examples{
# Build metadata for vectors
\dontrun{
-meta <- meta_for_ca_census_vectors("v_CA16_4836","v_CA16_4838","v_CA16_4899")
+meta <- meta_for_ca_census_vectors(c("v_CA16_4836","v_CA16_4838","v_CA16_4899"))
}
}
diff --git a/man/proportional_reaggregate.Rd b/man/proportional_reaggregate.Rd
index 45d81ad..4ab5caf 100644
--- a/man/proportional_reaggregate.Rd
+++ b/man/proportional_reaggregate.Rd
@@ -22,7 +22,7 @@ proportional_reaggregate(
\item{categories}{Vector of column names to re-aggregate}
\item{base}{Column name to use for proportional weighting when re-aggregating, or named vector with column name for each category.
-Categries that should be re-aggregated as means should be set to NA and will only be reaggregated if the base data has NA values.}
+Categories that should be re-aggregated as means should be set to NA and will only be reaggregated if the base data has NA values.}
}
\value{
dataframe with downsampled variables from parent_data
diff --git a/man/tongfen_aggregate.Rd b/man/tongfen_aggregate.Rd
index e54da9d..4e0db33 100644
--- a/man/tongfen_aggregate.Rd
+++ b/man/tongfen_aggregate.Rd
@@ -13,7 +13,10 @@ tongfen_aggregate(
)
}
\arguments{
-\item{data}{list of datasets to be aggregated}
+\item{data}{named list of datasets to be aggregated. The names identify the datasets, they are
+matched against the `geo_dataset` column in `meta` to pick the aggregation rules and labels
+for each dataset. Without names, or with names not found in `meta`, the rules for all
+datasets are applied and the variables keep their original names}
\item{correspondence}{correspondence data for gluing up the datasets}
@@ -32,19 +35,19 @@ aggregated dataset of class sf if base_geo is not NULL and data is of type sf or
\description{
\lifecycle{maturing}
-Aggregate variables secified in meta for several datasets according to correspondence.
+Aggregate variables specified in meta for several datasets according to correspondence.
}
\examples{
-# aggregate census tract level 2006 population data on common gepgraphy build through
+# aggregate census tract level 2006 and 2016 population data on common geography built through
# correspondence from 2006 and 2016 census tracts in the City of Vancouver.
\dontrun{
regions <- list(CSD="5915022")
geo1 <- cancensus::get_census("CA06",regions=regions,geo_format='sf',level='CT')
geo2 <- cancensus::get_census("CA16",regions=regions,geo_format='sf',level='CT')
-meta <- meta_for_additive_variables("CA06","Population")
+meta <- meta_for_additive_variables(c("CA06","CA16"),"Population")
correspondence <- get_tongfen_correspondence_ca_census(geo_datasets=c('CA06','CA16'),
regions=regions,level='CT')
-result <- tongfen_aggregate(list(geo1 \%>\% rename(GeoUIDCA06=GeoUID),
- geo2 \%>\% rename(GeoUIDCA16=GeoUID)),correspondence,meta)
+result <- tongfen_aggregate(list(CA06=geo1 \%>\% rename(GeoUIDCA06=GeoUID),
+ CA16=geo2 \%>\% rename(GeoUIDCA16=GeoUID)),correspondence,meta)
}
}
diff --git a/man/tongfen_anomaly_joins.Rd b/man/tongfen_anomaly_joins.Rd
new file mode 100644
index 0000000..10fdc63
--- /dev/null
+++ b/man/tongfen_anomaly_joins.Rd
@@ -0,0 +1,93 @@
+% Generated by roxygen2: do not edit by hand
+% Please edit documentation in R/tongfen_anomalies.R
+\name{tongfen_anomaly_joins}
+\alias{tongfen_anomaly_joins}
+\title{Determine regions to join to correct for likely geocoding anomalies}
+\usage{
+tongfen_anomaly_joins(
+ data,
+ variables,
+ id = "TongfenID",
+ neighbours = NULL,
+ rel_scale = 0.25,
+ abs_scale = 200,
+ p = 4,
+ surprise_cutoff = 0.15,
+ total_surprise_cutoff = 0.75,
+ cutoff_fact = 0.6,
+ surprise_reduction_const = 0.15,
+ sum_fact = 0.7
+)
+}
+\arguments{
+\item{data}{data on a common geography, with one row per region, for example as returned by
+`tongfen_aggregate` or `get_tongfen_ca_census`. Needs to be of class sf unless `neighbours` is specified.}
+
+\item{variables}{names of the columns holding the timeline of a count variable like population or
+dwellings, in temporal order. Changes from or to a missing value are not surprising and don't make up for
+surprising changes in neighbouring regions, and joined regions are missing a value if one of the regions they
+are made up of is. Replace missing values by zero beforehand if they stand for regions where nothing got counted}
+
+\item{id}{name of the column that uniquely identifies the regions, default is "TongfenID"}
+
+\item{neighbours}{optional, neighbouring regions as a table with the identifiers of pairs of neighbouring
+regions in the first two columns, or as a neighbours list like the ones returned by `spdep::poly2nb`.
+By default all regions with intersecting geometries are neighbours, which can miss neighbours
+if the geometries have been simplified and don't share their boundaries any more.}
+
+\item{rel_scale}{relative decrease that is half way to full surprise, default is `0.25` for a 25\% drop}
+
+\item{abs_scale}{absolute decrease that is half way to full surprise, default is `200`}
+
+\item{p}{exponent of the norm used to combine the surprises across the timeline into the total surprise.
+Large values focus on the most surprising change, 1 adds up the surprises of all changes, default is `4`}
+
+\item{surprise_cutoff}{changes with larger surprise count as surprising, default is `0.15`. Only
+regions with at least one surprising change are candidates}
+
+\item{total_surprise_cutoff}{only regions with larger total surprise are candidates, default is `0.75`}
+
+\item{cutoff_fact}{join regions if the total surprise after joining is lower than this share of the
+total surprise of the candidate region, default is `0.6`}
+
+\item{surprise_reduction_const}{join regions if joining lowers the total surprise by more than this,
+default is `0.15`}
+
+\item{sum_fact}{join regions if the total surprise after joining is lower than `cutoff_fact * sum_fact`
+times the sum of the total surprises of both regions, default is `0.7`}
+}
+\value{
+A tibble with one row for each region that gets joined with other regions, with the identifier
+of the region, the identifier of the joined region it becomes part of in the column named like the
+identifier with suffix `_joined`, by default `TongfenID_joined`, and the `round` in which the region
+first got joined to another region. The identifier of a joined region is the smallest
+identifier of the regions it is made up of.
+}
+\description{
+\lifecycle{experimental}
+
+Looks for regions with surprising drops in the timeline of a count variable that are complemented by
+a neighbouring region, as explained in `tongfen_detect_anomalies`, and joins them. This gets
+repeated on the joined regions until there are no more regions left that qualify to get joined.
+In each round a region only gets joined with one other region, the most surprising regions go first.
+
+Joining regions trades geographic detail for timelines that are consistent over time. The parameters
+control how aggressively regions get joined and are best calibrated on the data at hand, erring on the side of
+joining too few regions risks keeping geocoding problems, erring on the other side risks removing real
+changes and needlessly coarsens the geography.
+
+The result can be used to join the regions via `tongfen_join_regions`, or to update a correspondence
+via `tongfen_join_correspondence`.
+}
+\examples{
+# Correct 2001 through 2021 dissemination area level population timelines in the
+# City of Vancouver for likely geocoding problems
+\dontrun{
+datasets <- c("CA01","CA06","CA11","CA16","CA21")
+meta <- meta_for_additive_variables(datasets,"Population")
+data <- get_tongfen_ca_census(regions=list(CSD="5915022"),meta=meta,level="DA",base_geo="CA21")
+
+joins <- tongfen_anomaly_joins(data,paste0("Population_",datasets))
+corrected_data <- tongfen_join_regions(data,joins,meta)
+}
+}
diff --git a/man/tongfen_ca_census_ct.Rd b/man/tongfen_ca_census_ct.Rd
index 9d0ef11..0475e70 100644
--- a/man/tongfen_ca_census_ct.Rd
+++ b/man/tongfen_ca_census_ct.Rd
@@ -13,9 +13,9 @@ tongfen_ca_census_ct(
)
}
\arguments{
-\item{data1}{cancensus CT level datatset for year1 < year2 to serve as base for common geography}
+\item{data1}{cancensus CT level dataset for year1 < year2 to serve as base for common geography}
-\item{data2}{cancensus CT level datatset for year2 to be aggregated to common geography}
+\item{data2}{cancensus CT level dataset for year2 to be aggregated to common geography}
\item{data2_sum_vars}{vector of variable names to by summed up when aggregating geographies}
@@ -23,6 +23,10 @@ tongfen_ca_census_ct(
\item{na.rm}{optional parameter to remove NA values when summing, default = `TRUE`}
}
+\value{
+`data2` with the variables in `data2_sum_vars` aggregated to a common geography matching `data1`,
+identified by the `GeoUID` of `data1`
+}
\description{
\lifecycle{deprecated}
diff --git a/man/tongfen_detect_anomalies.Rd b/man/tongfen_detect_anomalies.Rd
new file mode 100644
index 0000000..f1b6758
--- /dev/null
+++ b/man/tongfen_detect_anomalies.Rd
@@ -0,0 +1,102 @@
+% Generated by roxygen2: do not edit by hand
+% Please edit documentation in R/tongfen_anomalies.R
+\name{tongfen_detect_anomalies}
+\alias{tongfen_detect_anomalies}
+\title{Detect likely geocoding anomalies in timelines on a common geography}
+\usage{
+tongfen_detect_anomalies(
+ data,
+ variables,
+ id = "TongfenID",
+ neighbours = NULL,
+ rel_scale = 0.25,
+ abs_scale = 200,
+ p = 4,
+ surprise_cutoff = 0.15,
+ total_surprise_cutoff = 0.75,
+ cutoff_fact = 0.6,
+ surprise_reduction_const = 0.15,
+ sum_fact = 0.7
+)
+}
+\arguments{
+\item{data}{data on a common geography, with one row per region, for example as returned by
+`tongfen_aggregate` or `get_tongfen_ca_census`. Needs to be of class sf unless `neighbours` is specified.}
+
+\item{variables}{names of the columns holding the timeline of a count variable like population or
+dwellings, in temporal order. Changes from or to a missing value are not surprising and don't make up for
+surprising changes in neighbouring regions, and joined regions are missing a value if one of the regions they
+are made up of is. Replace missing values by zero beforehand if they stand for regions where nothing got counted}
+
+\item{id}{name of the column that uniquely identifies the regions, default is "TongfenID"}
+
+\item{neighbours}{optional, neighbouring regions as a table with the identifiers of pairs of neighbouring
+regions in the first two columns, or as a neighbours list like the ones returned by `spdep::poly2nb`.
+By default all regions with intersecting geometries are neighbours, which can miss neighbours
+if the geometries have been simplified and don't share their boundaries any more.}
+
+\item{rel_scale}{relative decrease that is half way to full surprise, default is `0.25` for a 25\% drop}
+
+\item{abs_scale}{absolute decrease that is half way to full surprise, default is `200`}
+
+\item{p}{exponent of the norm used to combine the surprises across the timeline into the total surprise.
+Large values focus on the most surprising change, 1 adds up the surprises of all changes, default is `4`}
+
+\item{surprise_cutoff}{changes with larger surprise count as surprising, default is `0.15`. Only
+regions with at least one surprising change are candidates}
+
+\item{total_surprise_cutoff}{only regions with larger total surprise are candidates, default is `0.75`}
+
+\item{cutoff_fact}{join regions if the total surprise after joining is lower than this share of the
+total surprise of the candidate region, default is `0.6`}
+
+\item{surprise_reduction_const}{join regions if joining lowers the total surprise by more than this,
+default is `0.15`}
+
+\item{sum_fact}{join regions if the total surprise after joining is lower than `cutoff_fact * sum_fact`
+times the sum of the total surprises of both regions, default is `0.7`}
+}
+\value{
+A tibble with one row for each candidate region, most surprising first, with the identifier of the
+region, the number of surprising changes `surprise_count`, the total surprise `surprise_total`,
+the `period` with the most surprising change, the identifier of the `neighbour` that takes away most of
+the surprise, the total surprise `surprise_total_joined` after joining both and `join` indicating if both
+regions qualify to get joined.
+}
+\description{
+\lifecycle{experimental}
+
+TongFen is only as good as the geocoding that assigned the underlying data to geographic regions
+in the first place. Geocoding varies over time, and the same dwelling units, and the people living
+in them, can get assigned to different neighbouring regions in different years. In a timeline
+on a common geography this shows up as a surprising drop in one region that is offset by
+a corresponding jump in a neighbouring region.
+
+This function lists the candidate regions with surprising drops in the given count variable,
+together with the neighbouring region that takes away most of the surprise when both are joined.
+Use it to check for possible problems and to calibrate the parameters before joining regions with
+`tongfen_anomaly_joins`. Not all surprising drops are due to geocoding problems, a drop that is
+not complemented by a neighbouring region is likely real.
+
+The surprise of a change between two consecutive years ranges from 0 to 1. Only decreases are surprising,
+the surprise is the product of the surprise of the relative and of the absolute decrease, so that it takes
+a decrease that is large in both relative and absolute terms to be surprising. The total surprise of a
+region is the `p`-norm of the surprises across all changes in the timeline.
+
+A candidate region and its neighbour are flagged for joining if joining reduces the total surprise of the
+candidate region to below `cutoff_fact` times its total surprise, or reduces it by more than
+`surprise_reduction_const`, or reduces it to below `cutoff_fact * sum_fact` times the sum of
+the total surprises of both regions. To keep this comparable the surprise after joining is computed
+from the change of the joined regions relative to the counts of the candidate region alone.
+}
+\examples{
+# Check 2001 through 2021 dissemination area level population timelines in the
+# City of Vancouver for possible geocoding problems
+\dontrun{
+datasets <- c("CA01","CA06","CA11","CA16","CA21")
+meta <- meta_for_additive_variables(datasets,"Population")
+data <- get_tongfen_ca_census(regions=list(CSD="5915022"),meta=meta,level="DA",base_geo="CA21")
+
+anomalies <- tongfen_detect_anomalies(data,paste0("Population_",datasets))
+}
+}
diff --git a/man/tongfen_estimate.Rd b/man/tongfen_estimate.Rd
index 77604ef..311f114 100644
--- a/man/tongfen_estimate.Rd
+++ b/man/tongfen_estimate.Rd
@@ -17,7 +17,9 @@ on how to construct metadata.}
\item{na.rm}{remove NA values when aggregating, default is FALSE}
}
\value{
-`target` with estimated quantities from `source` as specified by `meta`
+`target` with estimated quantities from `source` as specified by `meta`, regions in `target`
+that don't overlap with `source` have `NA` values. Columns in `target` can't have the same name
+as the variables to be estimated.
}
\description{
\lifecycle{maturing}
diff --git a/man/tongfen_estimate_ca_census.Rd b/man/tongfen_estimate_ca_census.Rd
index 6d2e5f7..0f70149 100644
--- a/man/tongfen_estimate_ca_census.Rd
+++ b/man/tongfen_estimate_ca_census.Rd
@@ -39,6 +39,9 @@ one of these for all variables.}
\item{quiet}{suppress progress messages}
}
+\value{
+`geometry` with the estimated values for the census variables specified by `meta`
+}
\description{
\lifecycle{maturing}
@@ -48,8 +51,8 @@ optionally with downsampling via `proportional_reaggregate`,
to streamline estimating Canadian census data on custom geographies.
}
\examples{
-# Estimate a common geography for 2006 and 2016 dissemination areas in the City of Vancouver
-# based on the geographic data and check estimation errors
+# Estimate the 2016 population within 1 km of Toronto City Hall from dissemination area level
+# census data
\dontrun{
toronto_city_hall <- sf::st_point(c(-79.3839,43.6534)) \%>\%
sf::st_sfc(crs=4326) \%>\%
@@ -62,7 +65,7 @@ meta <- meta_for_additive_variables("CA16","Population")
data <- tongfen_estimate_ca_census(toronto_city_hall,meta,level="DA",intersection_level="CT")
print(paste0("Approximately ",scales::comma(data$Population,accuracy=100),
- " people live within a 1 km radius of Toronto City."))
+ " people live within a 1 km radius of Toronto City Hall."))
}
}
diff --git a/man/tongfen_join_correspondence.Rd b/man/tongfen_join_correspondence.Rd
new file mode 100644
index 0000000..76c8634
--- /dev/null
+++ b/man/tongfen_join_correspondence.Rd
@@ -0,0 +1,43 @@
+% Generated by roxygen2: do not edit by hand
+% Please edit documentation in R/tongfen_anomalies.R
+\name{tongfen_join_correspondence}
+\alias{tongfen_join_correspondence}
+\title{Join regions in a correspondence}
+\usage{
+tongfen_join_correspondence(correspondence, joins)
+}
+\arguments{
+\item{correspondence}{correspondence table with columns the unique geographic identifiers for each of the
+geographies and the TongfenID and TongfenUID, as for example returned by `estimate_tongfen_correspondence`
+or `get_tongfen_correspondence_ca_census`}
+
+\item{joins}{table with the regions to join as returned by `tongfen_anomaly_joins`, with columns
+`TongfenID` and `TongfenID_joined`}
+}
+\value{
+The correspondence with updated TongfenID and TongfenUID for the regions that got joined. If the
+correspondence has a TongfenMethod column "anomaly" gets added to the method of the regions that got joined.
+}
+\description{
+\lifecycle{experimental}
+
+Updates a correspondence so that the given regions are joined, for example to correct for likely geocoding
+anomalies as determined by `tongfen_anomaly_joins`. The updated correspondence can be used in `tongfen_aggregate`
+to aggregate data on the coarser common geography, which works for all variables `tongfen_aggregate` can
+deal with and for data that was not part of detecting the anomalies.
+}
+\examples{
+# Correct for likely geocoding problems in dissemination area level population timelines
+# and use the updated correspondence to aggregate data on the corrected common geography
+\dontrun{
+regions <- list(CSD="5915022")
+datasets <- c("CA01","CA06","CA11","CA16","CA21")
+meta <- meta_for_additive_variables(datasets,"Population")
+data <- get_tongfen_ca_census(regions=regions,meta=meta,level="DA",base_geo="CA21")
+joins <- tongfen_anomaly_joins(data,paste0("Population_",datasets))
+
+correspondence <- get_tongfen_correspondence_ca_census(geo_datasets=datasets,
+ regions=regions,level="DA") \%>\%
+ tongfen_join_correspondence(joins)
+}
+}
diff --git a/man/tongfen_join_regions.Rd b/man/tongfen_join_regions.Rd
new file mode 100644
index 0000000..759ba26
--- /dev/null
+++ b/man/tongfen_join_regions.Rd
@@ -0,0 +1,54 @@
+% Generated by roxygen2: do not edit by hand
+% Please edit documentation in R/tongfen_anomalies.R
+\name{tongfen_join_regions}
+\alias{tongfen_join_regions}
+\title{Join regions in data on a common geography}
+\usage{
+tongfen_join_regions(data, joins, meta = NULL, id = "TongfenID", na.rm = TRUE)
+}
+\arguments{
+\item{data}{data on a common geography, with one row per region, for example as returned by
+`tongfen_aggregate` or `get_tongfen_ca_census`}
+
+\item{joins}{table with the regions to join as returned by `tongfen_anomaly_joins`, with the identifier
+of the region and the identifier of the joined region it becomes part of in the column named like the
+identifier with suffix `_joined`}
+
+\item{meta}{optional metadata containing aggregation rules as for example returned by `meta_for_ca_census_vectors`,
+variables are matched by their label. Numeric variables that are not part of the metadata are treated as
+additive, if `NULL` (the default) that is the case for all numeric variables}
+
+\item{id}{name of the column that uniquely identifies the regions, default is "TongfenID"}
+
+\item{na.rm}{logical, determines how NA values should be treated when aggregating variables,
+default is `TRUE`}
+}
+\value{
+The data with the regions joined. Joined regions take the place and the identifier of the
+region with the smallest identifier among the regions they are made up of. Variables that are
+not numeric and not part of the metadata are `NA` for joined regions.
+}
+\description{
+\lifecycle{experimental}
+
+Joins regions in data that has already been aggregated to a common geography, for example to correct
+for likely geocoding anomalies as determined by `tongfen_anomaly_joins`. The data, and the geometries if
+the data is of class sf, of the regions that get joined are aggregated, all other regions are left as they are.
+
+Variables are aggregated according to the metadata, numeric variables that are not part of the metadata
+are assumed to be additive. Variables that are not additive, like averages, can only be aggregated if their
+parent variable is part of the data. If that is not the case use `tongfen_join_correspondence` to update the
+correspondence the data was built from and aggregate the original data again with `tongfen_aggregate`.
+}
+\examples{
+# Correct 2001 through 2021 dissemination area level population timelines in the
+# City of Vancouver for likely geocoding problems
+\dontrun{
+datasets <- c("CA01","CA06","CA11","CA16","CA21")
+meta <- meta_for_additive_variables(datasets,"Population")
+data <- get_tongfen_ca_census(regions=list(CSD="5915022"),meta=meta,level="DA",base_geo="CA21")
+
+joins <- tongfen_anomaly_joins(data,paste0("Population_",datasets))
+corrected_data <- tongfen_join_regions(data,joins,meta)
+}
+}
diff --git a/man/tongfen_tag_largest_overlap.Rd b/man/tongfen_tag_largest_overlap.Rd
index f002efc..ac9313a 100644
--- a/man/tongfen_tag_largest_overlap.Rd
+++ b/man/tongfen_tag_largest_overlap.Rd
@@ -14,8 +14,8 @@ tongfen_tag_largest_overlap(source, target, target_id)
\item{target_id}{name of the column in `target` table with unique id (character)}
}
\value{
-`source` with extra column with name `"target_id"` and column `...overlap_fraction` with
-the proportion of overlap of the target geometry with the respective `target_id`
+`source` with extra column with the name given by `target_id` and column `...overlap_fraction` with
+the proportion of the area of the source region that overlaps with the region in `target` with that id
}
\description{
\lifecycle{maturing}
@@ -23,11 +23,11 @@ the proportion of overlap of the target geometry with the respective `target_id`
tags regions in `source` by `target_id` of region in `target` with the largest overlap
}
\examples{
-# Estimate 2006 Populatino in the City of Vancouver dissemination ares on 2016 census geoographies
+# Tag 2016 dissemination areas in the City of Vancouver by the 2006 census tract they overlap
+# the most with
\dontrun{
-geo1 <- cancensus::get_census("CA06",regions=list(CSD="5915022"),geo_format='sf',level='DA')
+geo1 <- cancensus::get_census("CA06",regions=list(CSD="5915022"),geo_format='sf',level='CT')
geo2 <- cancensus::get_census("CA16",regions=list(CSD="5915022"),geo_format='sf',level='DA')
-meta <- meta_for_additive_variables("CA06","Population")
-result <- tongfen_estimate(geo2 \%>\% rename(Population_2016=Population),geo1,meta)
+result <- tongfen_tag_largest_overlap(geo2,geo1 \%>\% select(CT_2006=GeoUID),"CT_2006")
}
}
diff --git a/pkgdown/_pkgdown.yml b/pkgdown/_pkgdown.yml
index 950309d..54b1b72 100644
--- a/pkgdown/_pkgdown.yml
+++ b/pkgdown/_pkgdown.yml
@@ -25,6 +25,10 @@ articles:
navbar: ~
contents:
- tongfen_us
+ - title: "Geocoding anomalies"
+ navbar: ~
+ contents:
+ - tongfen_anomalies
authors:
Jens von Bergmann:
diff --git a/tests/testthat/test-aggregate.R b/tests/testthat/test-aggregate.R
index 378192e..6934ce2 100644
--- a/tests/testthat/test-aggregate.R
+++ b/tests/testthat/test-aggregate.R
@@ -234,3 +234,165 @@ test_that("tongfen_aggregate: returns a plain tibble when no dataset has geometr
result <- tongfen_aggregate(d$data, d$correspondence, meta)
expect_false("sf" %in% class(result))
})
+
+# ── same variable name in several datasets ───────────────────────────────────
+
+make_average_meta <- function(ds) {
+ bind_rows(
+ make_meta("hh", "Additive"),
+ make_meta("avg", "Average", parent = "hh")
+ ) %>%
+ mutate(dataset = ds, geo_dataset = ds,
+ label = paste0(.data$variable, "_", ds))
+}
+
+test_that("aggregate_data_with_meta: variables listed for several datasets are only scaled once", {
+ data <- tibble(
+ group = c("A", "A"),
+ hh = c(10, 30),
+ avg = c(2, 4)
+ ) %>% group_by(.data$group)
+
+ meta <- bind_rows(make_average_meta("X"), make_average_meta("Y"))
+
+ result <- aggregate_data_with_meta(data, meta, quiet = TRUE)
+ expect_equal(result$hh, 40)
+ expect_equal(result$avg, (2 * 10 + 4 * 30) / 40)
+})
+
+test_that("aggregate_data_with_meta: conflicting rules for the same variable are an error", {
+ data <- tibble(
+ group = c("A", "A"),
+ hh = c(10, 30),
+ avg = c(2, 4)
+ ) %>% group_by(.data$group)
+
+ meta <- bind_rows(
+ make_average_meta("X"),
+ make_meta("avg", "Additive") %>% mutate(geo_dataset = "Y")
+ )
+
+ expect_error(aggregate_data_with_meta(data, meta, quiet = TRUE), "Conflicting aggregation rules")
+})
+
+test_that("tongfen_aggregate: averages use the rules of their own dataset when variable names repeat", {
+ correspondence <- tibble(
+ GeoUIDX = c("x1", "x2"),
+ GeoUIDY = c("y1", "y1"),
+ TongfenID = c("x1", "x1"),
+ TongfenUID = c("u1", "u1")
+ )
+ data <- list(
+ X = tibble(GeoUIDX = c("x1", "x2"), hh = c(10, 30), avg = c(2, 4)),
+ Y = tibble(GeoUIDY = "y1", hh = 50, avg = 5)
+ )
+ meta <- bind_rows(make_average_meta("X"), make_average_meta("Y"))
+
+ result <- tongfen_aggregate(data, correspondence, meta)
+
+ expect_equal(nrow(result), 1L)
+ expect_equal(result$hh_X, 40)
+ expect_equal(result$avg_X, 3.5)
+ expect_equal(result$hh_Y, 50)
+ expect_equal(result$avg_Y, 5)
+
+ # same variable name but a different rule in the second dataset
+ meta_mixed <- bind_rows(
+ make_average_meta("X"),
+ make_average_meta("Y") %>% mutate(rule = "Additive", parent = NA_character_)
+ )
+ result <- tongfen_aggregate(data, correspondence, meta_mixed)
+ expect_equal(result$avg_X, 3.5)
+ expect_equal(result$avg_Y, 5)
+})
+
+# ── averages with missing values ─────────────────────────────────────────────
+
+test_that("aggregate_data_with_meta: missing averages don't count toward the weights with na.rm=TRUE", {
+ data <- tibble(
+ group = c("A", "A", "A", "B", "B"),
+ hh = c(10, 30, 60, 5, 15),
+ avg = c(2, NA, 4, NA, NA),
+ avg2 = c(1, 2, 3, 4, NA)
+ ) %>% group_by(.data$group)
+
+ meta <- bind_rows(
+ make_meta("hh", "Additive"),
+ make_meta("avg", "Average", parent = "hh"),
+ make_meta("avg2", "Average", parent = "hh")
+ )
+
+ result <- aggregate_data_with_meta(data, meta, na.rm = TRUE, quiet = TRUE)
+
+ expect_equal(names(result), c("group", "hh", "avg", "avg2"))
+ # the parent keeps counting all regions
+ expect_equal(result$hh, c(100, 20))
+ # the average is taken over the regions that have a value, each average on its own
+ expect_equal(result$avg[1], (2 * 10 + 4 * 60) / 70)
+ expect_true(is.na(result$avg[2]))
+ expect_equal(result$avg2, c((1 * 10 + 2 * 30 + 3 * 60) / 100, 4))
+
+ kept <- aggregate_data_with_meta(data, meta, na.rm = FALSE, quiet = TRUE)
+ expect_true(all(is.na(kept$avg)))
+ expect_equal(kept$avg2[1], (1 * 10 + 2 * 30 + 3 * 60) / 100)
+ expect_true(is.na(kept$avg2[2]))
+})
+
+# ── "Average to" variables ───────────────────────────────────────────────────
+
+make_average_to_meta <- function(variable, units = "Percentage (0-100)") {
+ make_meta(variable, "AverageTo", parent = "pop") %>% mutate(units = units)
+}
+
+test_that("aggregate_data_with_meta: percentage change is averaged relative to its base", {
+ # population grew from 100 to 110 and from 100 to 220
+ data <- tibble(
+ group = c("A", "A"),
+ pop = c(110, 220),
+ chg = c(10, 120)
+ ) %>% group_by(.data$group)
+
+ meta <- bind_rows(
+ make_meta("pop", "Additive") %>% mutate(units = "Number"),
+ make_average_to_meta("chg")
+ )
+
+ result <- aggregate_data_with_meta(data, meta, quiet = TRUE)
+ expect_equal(result$pop, 330)
+ expect_equal(result$chg, 65)
+ expect_equal(result$base_pop, 200)
+})
+
+test_that("aggregate_data_with_meta: 'Average to' variables sharing a parent don't interfere", {
+ data <- tibble(
+ group = c("A", "A"),
+ pop = c(110, 220),
+ chg1 = c(10, 10),
+ chg2 = c(100, 100),
+ ratio = c(0.1, 1.2)
+ ) %>% group_by(.data$group)
+
+ meta <- bind_rows(
+ make_meta("pop", "Additive") %>% mutate(units = "Number"),
+ make_average_to_meta("chg1"),
+ make_average_to_meta("chg2"),
+ make_average_to_meta("ratio", units = "Percentage ratio (0.0-1.0)")
+ )
+
+ result <- aggregate_data_with_meta(data, meta, quiet = TRUE)
+ expect_equal(result$chg1, 10)
+ expect_equal(result$chg2, 100)
+ expect_equal(result$ratio, 0.65)
+ # each variable comes with its own base
+ expect_equal(result$base_chg1, 300)
+ expect_equal(result$base_chg2, 165)
+ expect_equal(result$base_ratio, 200)
+
+ # same results as when aggregating the variables one at a time
+ for (v in c("chg1", "chg2", "ratio")) {
+ single <- aggregate_data_with_meta(data, meta %>% filter(.data$variable %in% c("pop", v)),
+ quiet = TRUE)
+ expect_equal(single[[v]], result[[v]])
+ expect_equal(single$base_pop, result[[paste0("base_", v)]])
+ }
+})
diff --git a/tests/testthat/test-anomalies.R b/tests/testthat/test-anomalies.R
new file mode 100644
index 0000000..80a1217
--- /dev/null
+++ b/tests/testthat/test-anomalies.R
@@ -0,0 +1,622 @@
+library(dplyr)
+library(sf)
+
+sq <- function(x0, y0, w, h) {
+ st_polygon(list(cbind(c(x0, x0 + w, x0 + w, x0, x0),
+ c(y0, y0, y0 + h, y0 + h, y0))))
+}
+
+# regions lined up in a row, each one only neighbours the one before and after
+row_regions <- function(ids, values) {
+ geometry <- st_sfc(lapply(seq_along(ids), \(i) sq(i - 1, 0, 1, 1)), crs = 3347)
+ st_sf(bind_cols(tibble(TongfenID = ids), as_tibble(values)), geometry = geometry)
+}
+
+years <- c("2001", "2006", "2011", "2016")
+
+# population moves from A to B between 2001 and 2006 while C stays flat, like dwellings
+# getting geocoded to the neighbouring region
+misallocation <- function() {
+ row_regions(c("A", "B", "C"),
+ rbind(c(1000, 400, 410, 420),
+ c(500, 1100, 1110, 1120),
+ c(800, 805, 810, 815)) %>%
+ `colnames<-`(years))
+}
+
+# population in A drops and no neighbour picks it up, like a redevelopment site getting cleared
+real_drop <- function() {
+ row_regions(c("A", "B", "C"),
+ rbind(c(700, 710, 720, 730),
+ c(1000, 300, 310, 320),
+ c(800, 805, 810, 815)) %>%
+ `colnames<-`(years))
+}
+
+# A loses to B between 2001 and 2006, B loses to C between 2006 and 2011. A only matches
+# up with its neighbour once B and C are joined.
+chain <- function() {
+ row_regions(c("A", "B", "C", "D"),
+ rbind(c(1000, 600, 600, 600),
+ c(500, 900, 400, 400),
+ c(500, 500, 1000, 1000),
+ c(800, 800, 800, 800)) %>%
+ `colnames<-`(years))
+}
+
+# ── surprise ──────────────────────────────────────────────────────────────────
+
+test_that("anomaly_surprise: only decreases that are large in relative and absolute terms surprise", {
+ surprise <- \(change, base) tongfen:::anomaly_surprise(change, base, rel_scale = 0.25, abs_scale = 200)
+
+ # half way to full surprise for both the relative and the absolute decrease
+ expect_equal(surprise(-200, 800), 0.25)
+ expect_equal(surprise(-400, 800), 0.75 * 0.75)
+ # large relative but small absolute decrease, and the other way around
+ expect_lt(surprise(-10, 20), 0.04)
+ expect_lt(surprise(-200, 100000), 0.01)
+ expect_equal(surprise(c(100, 0, NA, -50, 0), c(800, 800, 800, NA, 0)), rep(0, 5))
+ # decrease from nothing
+ expect_equal(surprise(-50, 0), 0)
+
+ m <- surprise(matrix(c(-200, 100, -400, NA), 2), matrix(800, 2, 2))
+ expect_equal(m, matrix(c(0.25, 0, 0.75 * 0.75, 0), 2))
+})
+
+# ── detection ─────────────────────────────────────────────────────────────────
+
+test_that("tongfen_detect_anomalies: complementary change in a neighbour qualifies for joining", {
+ anomalies <- tongfen_detect_anomalies(misallocation(), years, total_surprise_cutoff = 0.4)
+
+ expect_equal(names(anomalies),
+ c("TongfenID", "surprise_count", "surprise_total", "period", "neighbour",
+ "surprise_total_joined", "join"))
+ expect_equal(nrow(anomalies), 1L)
+ expect_equal(anomalies$TongfenID, "A")
+ expect_equal(anomalies$surprise_count, 1L)
+ expect_equal(anomalies$surprise_total, (1 - 0.5^(0.6 / 0.25)) * (1 - 0.5^3))
+ expect_equal(anomalies$period, "2001-2006")
+ expect_equal(anomalies$neighbour, "B")
+ expect_equal(anomalies$surprise_total_joined, 0)
+ expect_true(anomalies$join)
+})
+
+test_that("tongfen_detect_anomalies: drop without complementary change is listed but not joined", {
+ anomalies <- tongfen_detect_anomalies(real_drop(), years)
+
+ expect_equal(anomalies$TongfenID, "B")
+ expect_false(anomalies$join)
+ expect_equal(anomalies$surprise_total_joined, anomalies$surprise_total, tolerance = 0.01)
+ expect_equal(nrow(tongfen_anomaly_joins(real_drop(), years)), 0L)
+})
+
+test_that("tongfen_detect_anomalies: no candidates gives an empty table", {
+ data <- misallocation()
+ anomalies <- tongfen_detect_anomalies(data, years[2:4])
+ expect_equal(nrow(anomalies), 0L)
+ expect_equal(names(anomalies),
+ names(tongfen_detect_anomalies(data, years, total_surprise_cutoff = 0.4)))
+ expect_equal(nrow(tongfen_anomaly_joins(data, years[2:4])), 0L)
+})
+
+test_that("tongfen_detect_anomalies: candidates are ordered by surprise", {
+ anomalies <- tongfen_detect_anomalies(chain(), years, total_surprise_cutoff = 0.4)
+ expect_equal(anomalies$TongfenID, c("B", "A"))
+ expect_equal(anomalies$period, c("2006-2011", "2001-2006"))
+ expect_equal(anomalies$neighbour, c("C", "B"))
+ expect_equal(anomalies$join, c(TRUE, FALSE))
+})
+
+test_that("tongfen_detect_anomalies: missing values are not surprising", {
+ data <- misallocation()
+ data$`2006`[1] <- NA
+ expect_equal(nrow(tongfen_detect_anomalies(data, years, total_surprise_cutoff = 0.4)), 0L)
+ expect_equal(nrow(tongfen_anomaly_joins(data, years, total_surprise_cutoff = 0.4)), 0L)
+
+ # joined regions are missing the values their parts are missing
+ data <- chain()
+ data$`2016`[3] <- NA
+ joins <- tongfen_anomaly_joins(data, years, total_surprise_cutoff = 0.4)
+ expect_equal(joins$TongfenID, c("A", "B", "C"))
+ # a missing value in a neighbour does not make up for a surprising change
+ data$`2006`[3] <- NA
+ anomalies <- tongfen_detect_anomalies(data, years, total_surprise_cutoff = 0.4)
+ expect_equal(anomalies$TongfenID, c("B", "A"))
+ expect_equal(anomalies$surprise_total_joined[1], anomalies$surprise_total[1])
+ expect_equal(anomalies$join, c(FALSE, FALSE))
+ expect_equal(nrow(tongfen_anomaly_joins(data, years, total_surprise_cutoff = 0.4)), 0L)
+})
+
+test_that("tongfen_detect_anomalies: candidate without neighbours", {
+ data <- misallocation() %>% st_drop_geometry()
+ anomalies <- tongfen_detect_anomalies(data, years, total_surprise_cutoff = 0.4,
+ neighbours = tibble(a = "B", b = "C"))
+ expect_equal(anomalies$TongfenID, "A")
+ expect_true(is.na(anomalies$neighbour))
+ expect_true(is.na(anomalies$surprise_total_joined))
+ expect_false(anomalies$join)
+})
+
+# ── joins ─────────────────────────────────────────────────────────────────────
+
+test_that("tongfen_anomaly_joins: joins regions with complementary changes", {
+ joins <- tongfen_anomaly_joins(misallocation(), years, total_surprise_cutoff = 0.4)
+ expect_equal(joins, tibble(TongfenID = c("A", "B"), TongfenID_joined = "A", round = 1L))
+
+ # too high a bar for regions to become candidates
+ expect_equal(nrow(tongfen_anomaly_joins(misallocation(), years, total_surprise_cutoff = 0.9)), 0L)
+})
+
+test_that("tongfen_anomaly_joins: joined regions keep getting joined", {
+ joins <- tongfen_anomaly_joins(chain(), years, total_surprise_cutoff = 0.4)
+ expect_equal(joins, tibble(TongfenID = c("A", "B", "C"), TongfenID_joined = "A",
+ round = c(2L, 1L, 1L)))
+})
+
+test_that("tongfen_anomaly_joins: joined regions are named after their smallest identifier", {
+ data <- chain() %>% mutate(TongfenID = c("59150010", "5915002", "59150003", "59150004"))
+ joins <- tongfen_anomaly_joins(data, years, total_surprise_cutoff = 0.4)
+ # sorted like the identifiers of a correspondence, character by character
+ expect_equal(joins$TongfenID, c("59150003", "59150010", "5915002"))
+ expect_true(all(joins$TongfenID_joined == "59150003"))
+})
+
+test_that("tongfen_anomaly_joins: result does not depend on the order of the regions", {
+ # A drops and both neighbours make up for it in full
+ data <- row_regions(c("B", "A", "C"),
+ rbind(c(500, 1100, 1110, 1120),
+ c(1000, 400, 410, 420),
+ c(500, 1100, 1110, 1120)) %>%
+ `colnames<-`(years))
+ expected <- tibble(TongfenID = c("A", "B"), TongfenID_joined = "A", round = 1L)
+ expect_equal(tongfen_anomaly_joins(data, years, total_surprise_cutoff = 0.4), expected)
+ expect_equal(tongfen_anomaly_joins(data[3:1, ], years, total_surprise_cutoff = 0.4), expected)
+ expect_equal(tongfen_detect_anomalies(data[3:1, ], years, total_surprise_cutoff = 0.4),
+ tongfen_detect_anomalies(data, years, total_surprise_cutoff = 0.4))
+})
+
+test_that("tongfen_anomaly_joins: identifier column can go by any name and type", {
+ data <- chain() %>% st_drop_geometry() %>% mutate(TongfenID = c(40L, 30L, 20L, 10L)) %>%
+ rename(GeoUID = "TongfenID")
+ joins <- tongfen_anomaly_joins(data, years, id = "GeoUID", total_surprise_cutoff = 0.4,
+ neighbours = tibble(a = c(40L, 30L, 20L), b = c(30L, 20L, 10L)))
+ expect_equal(joins, tibble(GeoUID = c(20L, 30L, 40L), GeoUID_joined = 20L,
+ round = c(1L, 1L, 2L)))
+})
+
+test_that("tongfen_anomaly_joins: neighbours can be specified", {
+ data <- chain()
+ expected <- tongfen_anomaly_joins(data, years, total_surprise_cutoff = 0.4)
+ data <- data %>% st_drop_geometry()
+
+ neighbour_table <- tibble(id1 = c("A", "B", "C", "A", "X"), id2 = c("B", "C", "D", "A", "A"))
+ expect_equal(tongfen_anomaly_joins(data, years, total_surprise_cutoff = 0.4,
+ neighbours = neighbour_table),
+ expected)
+
+ # neighbours list in the format of the spdep package
+ nb <- list(2L, c(1L, 3L), c(2L, 4L), 3L)
+ attr(nb, "region.id") <- data$TongfenID
+ class(nb) <- "nb"
+ expect_equal(tongfen_anomaly_joins(data, years, total_surprise_cutoff = 0.4, neighbours = nb),
+ expected)
+
+ # in a different order than the data, and with a region without neighbours
+ nb <- list(0L, 3L, c(2L, 4L), c(3L, 5L), 4L)
+ attr(nb, "region.id") <- c("Z", rev(data$TongfenID))
+ expect_equal(tongfen_anomaly_joins(data, years, total_surprise_cutoff = 0.4, neighbours = nb),
+ expected)
+
+ # without the link between B and C nothing can get joined
+ expect_equal(nrow(tongfen_anomaly_joins(data, years, total_surprise_cutoff = 0.4,
+ neighbours = neighbour_table[-2, ])), 0L)
+})
+
+test_that("tongfen_anomaly_joins: checks inputs", {
+ data <- misallocation()
+ expect_error(tongfen_anomaly_joins(st_drop_geometry(data), years), "neighbours")
+ expect_error(tongfen_anomaly_joins(data, years, neighbours = "B"), "neighbours")
+ expect_error(tongfen_anomaly_joins(data, "2001"), "at least two")
+ expect_error(tongfen_anomaly_joins(data, c("2001", "2026")), "2026")
+ expect_error(tongfen_anomaly_joins(data, years, id = "GeoUID"), "GeoUID")
+ expect_error(tongfen_anomaly_joins(data %>% mutate(`2006` = as.character(.data$`2006`)), years),
+ "numeric")
+ expect_error(tongfen_anomaly_joins(data %>% mutate(TongfenID = c("A", "A", "B")), years),
+ "uniquely")
+ expect_error(tongfen_anomaly_joins(data, years, p = c(2, 4)), "single number")
+ expect_error(tongfen_detect_anomalies(data, years, abs_scale = 0), "positive")
+})
+
+# ── comparison with the original implementation ───────────────────────────────
+
+# The method was first implemented in
+# https://doodles.mountainmath.ca/posts/2024-07-26-geocoding-errors-in-aggregate-data/
+# This is the code from that post, with the neighbours derived via `st_intersects` instead
+# of `spdep::poly2nb`. It works on timelines in columns named by year and names joined
+# regions by stringing together their identifiers.
+blog_rel_surprise <- function(x) {
+ r <- as.integer(x <= 0 & !is.na(x) & is.finite(x))
+ r[r == 1] <- 1 - dexp(x[r == 1] * log(0.5) / 0.25, rate = 1)
+ r
+}
+
+blog_abs_surprise <- function(x) {
+ r <- as.integer(x <= 0 & !is.na(x) & is.finite(x))
+ r[r == 1] <- 1 - dexp(x[r == 1] * log(0.5) / 200, rate = 1)
+ r
+}
+
+blog_add_surprise <- function(data, base_field = "value") {
+ data |>
+ arrange(Year) |>
+ mutate(`Absolute change` = value - lag(value, order_by = Year),
+ `Relative change` = `Absolute change` / lag(!!as.name(base_field), order_by = Year),
+ .by = TongfenID) |>
+ mutate(surprise = (blog_rel_surprise(`Relative change`)) * (blog_abs_surprise(`Absolute change`)))
+}
+
+blog_get_match_list <- function(geo_data2,
+ cutoff_fact = 0.6, total_surprise_cutoff = 0.75,
+ surprise_reduction_const = 0.15, sum_fact = 0.7, p = 4) {
+ intersecting <- st_intersects(geo_data2)
+ data_nb <- lapply(seq_along(intersecting), \(i) geo_data2$TongfenID[setdiff(intersecting[[i]], i)]) |>
+ setNames(geo_data2$TongfenID)
+
+ summarize_surprise <- function(data, fields = "surprise") {
+ data |>
+ summarize(across(all_of(fields), list(count = ~sum(.x > 0.15),
+ total = ~(sum(.x^p))^(1 / p))),
+ .by = TongfenID) |>
+ arrange(!!as.name(paste0(fields[1], "_count")), !!as.name(paste0(fields[1], "_total")))
+ }
+
+ long_data2 <- geo_data2 |>
+ st_drop_geometry() |>
+ select(TongfenID, matches("^\\d{4}$")) |>
+ tidyr::pivot_longer(matches("^\\d{4}$"), names_to = "Year") |>
+ blog_add_surprise()
+
+ candidate_list <- long_data2 |>
+ summarize_surprise() |>
+ purrr::map_df(rev) |>
+ filter(surprise_count > 0, surprise_total > total_surprise_cutoff)
+
+ match_list <- tibble(TongfenID = NA_character_, TongfenID_original = NA_character_) |>
+ slice(-1)
+
+ if (nrow(candidate_list) == 0) return(match_list)
+
+ for (i in 1:nrow(candidate_list)) {
+ id0 <- candidate_list$TongfenID[i]
+
+ if (id0 %in% match_list$TongfenID_original) next
+
+ neighbour_ids <- data_nb[[id0]]
+
+ if (length(neighbour_ids) == 0) next
+
+ od <- long_data2 |>
+ filter(TongfenID %in% c(id0))
+
+ original <- od |>
+ summarize_surprise()
+
+ reductions <- long_data2 |>
+ filter(TongfenID %in% c(neighbour_ids)) |>
+ left_join(od |> select(Year, ov = value), by = "Year") |>
+ mutate(value = value + ov) |>
+ blog_add_surprise(base_field = "ov") |>
+ rename(surprise_ov = surprise) |>
+ blog_add_surprise() |>
+ summarize_surprise(c("surprise_ov", "surprise")) |>
+ slice(1) |>
+ rename(TongfenID2 = TongfenID) |>
+ mutate(TongfenID1 = original$TongfenID, .before = TongfenID2) |>
+ mutate(TongfenID = paste0(TongfenID1, "_", TongfenID2))
+
+ if (reductions$TongfenID2 %in% c(match_list$TongfenID_original)) next
+
+ pre_sum <- long_data2 |>
+ filter(TongfenID %in% c(reductions$TongfenID1, reductions$TongfenID2)) |>
+ summarize_surprise()
+
+ if (reductions$surprise_ov_total < cutoff_fact * original$surprise_total |
+ original$surprise_total - reductions$surprise_ov_total > surprise_reduction_const |
+ reductions$surprise_ov_total < cutoff_fact * sum_fact * sum(pre_sum$surprise_total)) {
+ match_list <- bind_rows(match_list,
+ reductions |>
+ tidyr::pivot_longer(c(TongfenID1, TongfenID2),
+ values_to = "TongfenID_original"))
+ }
+ }
+ match_list
+}
+
+blog_iterate_geo_match_joins <- function(geo_data2, ...) {
+ stop_looking <- FALSE
+
+ while (!stop_looking) {
+ match_list <- blog_get_match_list(geo_data2, ...)
+
+ if (nrow(match_list) == 0) {
+ stop_looking <- TRUE
+ next
+ }
+
+ g1 <- geo_data2 |>
+ rename(TongfenID_original = TongfenID) |>
+ inner_join(match_list |> select(TongfenID, TongfenID_original),
+ by = "TongfenID_original",
+ relationship = "many-to-one") |>
+ mutate(TongfenID = coalesce(TongfenID, TongfenID_original)) |>
+ select(-TongfenID_original) |>
+ group_by(TongfenID) |>
+ summarize(across(matches("\\d{4}"), sum), .groups = "drop") |>
+ st_make_valid()
+
+ g2 <- geo_data2 |>
+ filter(!(TongfenID %in% match_list$TongfenID_original))
+
+ geo_data2 <- bind_rows(g1, g2)
+ }
+ geo_data2
+}
+
+# Regions on a grid with slowly declining counts, and counts that get moved between
+# neighbouring regions for a couple of years. The counts are not rounded and all
+# regions change in all periods to keep clear of ties, the two implementations
+# order regions differently after the first round of joins.
+random_timelines <- function(n, n_moves, n_drops, seed) {
+ set.seed(seed)
+ cells <- expand.grid(x = seq_len(n), y = seq_len(n))
+ n_years <- 6
+ values <- matrix(runif(n * n, 300, 1500), n * n, n_years) -
+ t(apply(matrix(runif(n * n * n_years, 1, 20), n * n, n_years), 1, cumsum))
+ cell_index <- \(x, y) (y - 1) * n + x
+
+ for (move in seq_len(n_moves)) {
+ from <- sample(n * n, 1)
+ step <- list(c(1, 0), c(-1, 0), c(0, 1), c(0, -1))[[sample(4, 1)]]
+ x <- cells$x[from] + step[1]
+ y <- cells$y[from] + step[2]
+ if (x < 1 || x > n || y < 1 || y > n) next
+ to <- cell_index(x, y)
+ start <- sample(2:n_years, 1)
+ moved_years <- start:sample(start:n_years, 1)
+ amount <- runif(1, 0.2, 0.8) * min(values[from, ])
+ values[from, moved_years] <- values[from, moved_years] - amount
+ values[to, moved_years] <- values[to, moved_years] + amount
+ }
+ for (drop in seq_len(n_drops)) {
+ region <- sample(n * n, 1)
+ dropped_years <- sample(2:n_years, 1):n_years
+ values[region, dropped_years] <- values[region, dropped_years] * runif(1, 0.3, 0.7)
+ }
+ colnames(values) <- seq(1996, by = 5, length.out = n_years)
+
+ geometry <- st_sfc(lapply(seq_len(n * n), \(i) sq(cells$x[i], cells$y[i], 1, 1)), crs = 3347)
+ st_sf(bind_cols(tibble(TongfenID = sprintf("r%03d", sample(n * n))), as_tibble(values)),
+ geometry = geometry)
+}
+
+test_that("tongfen_anomaly_joins: agrees with the original implementation", {
+ rounds <- c()
+ for (seed in 1:4) {
+ data <- random_timelines(n = 7, n_moves = 30, n_drops = 5, seed = seed)
+ timeline <- names(data)[grepl("^\\d{4}$", names(data))]
+
+ joins <- tongfen_anomaly_joins(data, timeline, total_surprise_cutoff = 0.4)
+ blog_ids <- blog_iterate_geo_match_joins(data, total_surprise_cutoff = 0.4)$TongfenID
+ blog_groups <- strsplit(blog_ids, "_", fixed = TRUE)
+ blog_groups <- blog_groups[lengths(blog_groups) > 1] %>%
+ vapply(\(g) paste0(sort(g), collapse = "_"), character(1))
+ groups <- split(joins$TongfenID, joins$TongfenID_joined) %>%
+ vapply(\(g) paste0(sort(g), collapse = "_"), character(1))
+
+ expect_gt(length(groups), 3)
+ expect_setequal(unname(groups), blog_groups)
+ rounds <- c(rounds, max(joins$round))
+
+ # stricter parameters
+ joins <- tongfen_anomaly_joins(data, timeline, cutoff_fact = 0.3, surprise_reduction_const = 0.4,
+ sum_fact = 0.5, p = 2)
+ blog_ids <- blog_iterate_geo_match_joins(data, cutoff_fact = 0.3, surprise_reduction_const = 0.4,
+ sum_fact = 0.5, p = 2)$TongfenID
+ expect_equal(nrow(data) - n_distinct(joins$TongfenID) + n_distinct(joins$TongfenID_joined),
+ length(blog_ids))
+ expect_setequal(joins$TongfenID, unlist(strsplit(blog_ids[grepl("_", blog_ids)], "_", fixed = TRUE)))
+ }
+ # some of the timelines need several rounds of joins
+ expect_gt(max(rounds), 1)
+})
+
+test_that("tongfen_anomaly_joins: random timelines don't depend on the order of the regions", {
+ data <- random_timelines(n = 6, n_moves = 25, n_drops = 4, seed = 11)
+ timeline <- names(data)[grepl("^\\d{4}$", names(data))]
+ joins <- tongfen_anomaly_joins(data, timeline, total_surprise_cutoff = 0.4)
+ expect_gt(nrow(joins), 6)
+ expect_equal(tongfen_anomaly_joins(data[sample(nrow(data)), ], timeline, total_surprise_cutoff = 0.4),
+ joins)
+})
+
+# ── joining regions ───────────────────────────────────────────────────────────
+
+joins_ab <- tibble(TongfenID = c("A", "B"), TongfenID_joined = "A", round = 1L)
+
+test_that("tongfen_join_regions: aggregates joined regions and leaves the rest alone", {
+ data <- misallocation() %>%
+ mutate(TongfenUID = paste0("GeoUID16:", c("1", "2,3", "4"), " GeoUID21:", c("11,12", "13", "14")),
+ .after = "TongfenID")
+
+ expect_message(result <- tongfen_join_regions(data, joins_ab), "additive")
+ expect_s3_class(result, "sf")
+ expect_equal(names(result), names(data))
+ expect_equal(result$TongfenID, c("A", "C"))
+ expect_equal(result$TongfenUID, c("GeoUID16:1,2,3 GeoUID21:11,12,13", "GeoUID16:4 GeoUID21:14"))
+ expect_equal(result$`2001`, c(1500, 800))
+ expect_equal(result$`2016`, c(1540, 815))
+ expect_equal(as.numeric(st_area(result)), c(2, 1))
+ expect_true(st_equals(st_geometry(result)[1], sq(0, 0, 2, 1), sparse = FALSE)[1, 1])
+ expect_true(st_equals(st_geometry(result)[2], st_geometry(data)[3], sparse = FALSE)[1, 1])
+ expect_equal(st_crs(result), st_crs(data))
+ expect_equal(result %>% st_drop_geometry() %>% filter(.data$TongfenID == "C"),
+ data %>% st_drop_geometry() %>% filter(.data$TongfenID == "C"))
+
+ # the timeline of the joined region does not have anything surprising left
+ expect_equal(nrow(tongfen_detect_anomalies(result, years, total_surprise_cutoff = 0.4)), 0L)
+})
+
+test_that("tongfen_join_regions: works on data without geometry and keeps the order of regions", {
+ data <- chain() %>% st_drop_geometry() %>% slice(c(4, 2, 3, 1))
+ joins <- tibble(TongfenID = c("C", "B"), TongfenID_joined = "B")
+
+ result <- suppressMessages(tongfen_join_regions(data, joins))
+ expect_false(inherits(result, "sf"))
+ expect_equal(result$TongfenID, c("D", "B", "A"))
+ expect_equal(result$`2011`, c(800, 1400, 600))
+
+ # nothing to join
+ expect_equal(tongfen_join_regions(data, joins[0, ]), data)
+ expect_equal(tongfen_join_regions(data, tibble(TongfenID = "X", TongfenID_joined = "Y")), data)
+ # regions other regions get joined to don't have to be listed
+ result <- suppressMessages(tongfen_join_regions(data, joins[1, ]))
+ expect_equal(result$TongfenID, c("D", "B", "A"))
+ expect_equal(result$`2011`, c(800, 1400, 600))
+})
+
+test_that("tongfen_join_regions: aggregates according to metadata", {
+ data <- tibble(TongfenID = c("A", "B", "C"),
+ Population_CA16 = c(100, 300, 50),
+ Income_CA16 = c(10, 20, 30),
+ Population_CA21 = c(300, 100, 60),
+ Income_CA21 = c(40, 20, 35),
+ name = c("a", "b", "c"))
+ meta <- tibble(variable = c("v_pop", "v_income", "v_pop", "v_income"),
+ label = c("Population_CA16", "Income_CA16", "Population_CA21", "Income_CA21"),
+ rule = c("Additive", "Average", "Additive", "Average"),
+ parent = c(NA, "v_pop", NA, "v_pop"),
+ type = "Original",
+ geo_dataset = c("CA16", "CA16", "CA21", "CA21"))
+
+ expect_message(result <- tongfen_join_regions(data, joins_ab, meta), "name")
+ expect_equal(result$TongfenID, c("A", "C"))
+ expect_equal(result$Population_CA16, c(400, 50))
+ expect_equal(result$Income_CA16, c((100 * 10 + 300 * 20) / 400, 30))
+ expect_equal(result$Income_CA21, c((300 * 40 + 100 * 20) / 400, 35))
+ expect_equal(result$name, c(NA, "c"))
+
+ # count variables that are not part of the metadata, like the ones census calls add
+ expect_message(result <- tongfen_join_regions(data %>% mutate(Dwellings_CA16 = c(40, 110, 20)),
+ joins_ab, meta),
+ "not part of the metadata as additive: Dwellings_CA16")
+ expect_equal(result$Dwellings_CA16, c(150, 20))
+ expect_equal(result$Income_CA16, c((100 * 10 + 300 * 20) / 400, 30))
+
+ # averages can't be aggregated without their parent
+ expect_error(tongfen_join_regions(data %>% select(-"Population_CA21"), joins_ab, meta),
+ "Income_CA21.*tongfen_join_correspondence")
+ # treating averages as additive is on the user
+ result <- suppressMessages(tongfen_join_regions(data, joins_ab))
+ expect_equal(result$Income_CA16, c(30, 30))
+})
+
+test_that("tongfen_join_regions: checks inputs", {
+ data <- misallocation()
+ expect_error(tongfen_join_regions(data, joins_ab %>% select(-"TongfenID_joined")), "TongfenID_joined")
+ expect_error(tongfen_join_regions(data, joins_ab, id = "GeoUID"), "GeoUID")
+ expect_error(tongfen_join_regions(bind_rows(data, data), joins_ab), "uniquely")
+})
+
+test_that("merge_tongfen_uids: combines the identifiers by geography", {
+ merge_uids <- tongfen:::merge_tongfen_uids
+ expect_equal(merge_uids(c("a:2,3 b:10", "a:1 b:9,10")), "a:1,2,3 b:10,9")
+ expect_equal(merge_uids(c("a:1 b:9,10", "a:2,3 b:10")), "a:1,2,3 b:10,9")
+ expect_equal(merge_uids("a:1 b:9"), "a:1 b:9")
+ # regions that are not part of all geographies
+ expect_equal(merge_uids(c("a:1 ", "a:2 b:3")), "a:1,2 b:3")
+ # not in the format of a TongfenUID
+ expect_equal(merge_uids(c("x", "y", "x")), "x y")
+})
+
+# ── joining correspondences ───────────────────────────────────────────────────
+
+# two geographies on a grid of four regions, the second one has the top row in one piece
+two_geographies <- function() {
+ ids <- c("a1", "a2", "a3", "a4")
+ geo_a <- st_sf(idA = ids, Population = c(1000, 400, 300, 500),
+ geometry = st_sfc(sq(0, 0, 1, 1), sq(1, 0, 1, 1), sq(0, 1, 1, 1), sq(1, 1, 1, 1),
+ crs = 3347))
+ geo_b <- st_sf(idB = c("b1", "b2", "b3"), Population = c(400, 1050, 820),
+ geometry = st_sfc(sq(0, 0, 1, 1), sq(1, 0, 1, 1), sq(0, 1, 2, 1), crs = 3347))
+ correspondence <- tibble(idA = ids, idB = c("b1", "b2", "b3", "b3")) %>%
+ tongfen:::get_tongfen_correspondence()
+ list(data = list(A = geo_a, B = geo_b), correspondence = correspondence,
+ meta = meta_for_additive_variables(c("A", "B"), "Population"))
+}
+
+test_that("tongfen_join_correspondence: joins the regions in the correspondence", {
+ correspondence <- two_geographies()$correspondence
+ joins <- tibble(TongfenID = c("a1", "a2"), TongfenID_joined = "a1")
+
+ result <- tongfen_join_correspondence(correspondence, joins)
+ expect_equal(names(result), names(correspondence))
+ expect_equal(result$TongfenID, c("a1", "a1", "a3", "a3"))
+ expect_equal(result$TongfenUID, c(rep("idA:a1,a2 idB:b1,b2", 2), rep("idA:a3,a4 idB:b3", 2)))
+ expect_equal(result[3:4, ], correspondence[3:4, ])
+
+ # joining a region made up of several regions
+ result <- tongfen_join_correspondence(correspondence, tibble(TongfenID = c("a3", "a2"),
+ TongfenID_joined = "a2"))
+ expect_equal(result$TongfenID, c("a1", "a2", "a2", "a2"))
+ expect_equal(result$TongfenUID, c("idA:a1 idB:b1", rep("idA:a2,a3,a4 idB:b2,b3", 3)))
+
+ expect_equal(tongfen_join_correspondence(correspondence, joins[0, ]), correspondence)
+ expect_warning(result <- tongfen_join_correspondence(correspondence,
+ bind_rows(joins, tibble(TongfenID = "x",
+ TongfenID_joined = "a1"))),
+ "Did not find 1")
+ expect_equal(result$TongfenID, c("a1", "a1", "a3", "a3"))
+ expect_error(tongfen_join_correspondence(correspondence, tibble(TongfenID = "a1")),
+ "TongfenID_joined")
+})
+
+test_that("tongfen_join_correspondence: tags the method of joined regions", {
+ correspondence <- two_geographies()$correspondence %>%
+ mutate(TongfenMethod = c("identifier", "identifier", "estimate", "estimate"))
+ joins <- tibble(TongfenID = c("a1", "a2"), TongfenID_joined = "a1")
+
+ result <- tongfen_join_correspondence(correspondence, joins)
+ expect_equal(result$TongfenMethod,
+ c("identifier, anomaly", "identifier, anomaly", "estimate", "estimate"))
+ # joining again does not tag again
+ joins <- tibble(TongfenID = c("a1", "a3"), TongfenID_joined = "a1")
+ result <- tongfen_join_correspondence(result, joins)
+ expect_equal(result$TongfenID, rep("a1", 4))
+ expect_equal(result$TongfenMethod,
+ c("identifier, anomaly", "identifier, anomaly", "estimate, anomaly", "estimate, anomaly"))
+ # the summary of the correspondence still works
+ expect_equal(nrow(check_tongfen_areas(two_geographies()$data, result)), 1L)
+})
+
+test_that("joining the aggregated data and aggregating with the joined correspondence agree", {
+ fixture <- two_geographies()
+ aggregated <- tongfen_aggregate(fixture$data, fixture$correspondence, fixture$meta, base_geo = "A")
+ expect_equal(aggregated$TongfenID, c("a1", "a2", "a3"))
+
+ # a1 loses 600 that show up in a2
+ joins <- tongfen_anomaly_joins(aggregated, c("Population_A", "Population_B"),
+ total_surprise_cutoff = 0.4)
+ expect_equal(joins$TongfenID, c("a1", "a2"))
+
+ joined <- tongfen_join_regions(aggregated, joins, fixture$meta)
+ reaggregated <- tongfen_aggregate(fixture$data,
+ tongfen_join_correspondence(fixture$correspondence, joins),
+ fixture$meta, base_geo = "A")
+
+ expect_equal(joined$TongfenID, c("a1", "a3"))
+ expect_equal(joined$Population_B, c(1450, 820))
+ expect_equal(st_drop_geometry(joined)[names(st_drop_geometry(reaggregated))],
+ st_drop_geometry(reaggregated))
+ expect_true(all(diag(st_equals(joined, reaggregated, sparse = FALSE))))
+ expect_equal(as.character(st_geometry_type(joined)), as.character(st_geometry_type(reaggregated)))
+})
diff --git a/tests/testthat/test-cached-download.R b/tests/testthat/test-cached-download.R
new file mode 100644
index 0000000..abccf95
--- /dev/null
+++ b/tests/testthat/test-cached-download.R
@@ -0,0 +1,109 @@
+# ── cached_download ───────────────────────────────────────────────────────────
+
+# local file standing in for the remote file, with its md5 checksum as ETag like S3
+local_remote <- function(content) {
+ remote <- tempfile()
+ writeLines(content, remote)
+ url <- paste0("file://", normalizePath(remote, winslash = "/"))
+ rm(list = intersect(url, ls(tongfen:::tongfen_session)), envir = tongfen:::tongfen_session)
+ list(path = remote, url = url)
+}
+
+test_that("cached_download: downloads the file and remembers the ETag", {
+ remote <- local_remote("a")
+ path <- file.path(tempfile(), "file.txt")
+ local_mocked_bindings(remote_etag = function(url) unname(tools::md5sum(remote$path)))
+ expect_equal(tongfen:::cached_download(remote$url, path), path)
+ expect_equal(readLines(path), "a")
+ expect_equal(readLines(paste0(path, ".etag")), unname(tools::md5sum(remote$path)))
+})
+
+test_that("cached_download: checks the remote file only once per session", {
+ remote <- local_remote("a")
+ path <- file.path(tempfile(), "file.txt")
+ checks <- 0
+ local_mocked_bindings(remote_etag = function(url) {
+ checks <<- checks + 1
+ unname(tools::md5sum(remote$path))
+ })
+ tongfen:::cached_download(remote$url, path)
+ tongfen:::cached_download(remote$url, path)
+ expect_equal(checks, 1)
+ tongfen:::cached_download(remote$url, path, refresh = TRUE)
+ expect_equal(checks, 2)
+})
+
+test_that("cached_download: downloads again only if the remote file changed", {
+ remote <- local_remote("a")
+ path <- file.path(tempfile(), "file.txt")
+ local_mocked_bindings(remote_etag = function(url) unname(tools::md5sum(remote$path)))
+ tongfen:::cached_download(remote$url, path)
+
+ # new session, unchanged remote file, the cached file is left alone
+ rm(list = remote$url, envir = tongfen:::tongfen_session)
+ writeLines("local", path)
+ tongfen:::cached_download(remote$url, path)
+ expect_equal(readLines(path), "local")
+
+ # new session, changed remote file
+ rm(list = remote$url, envir = tongfen:::tongfen_session)
+ writeLines("b", remote$path)
+ tongfen:::cached_download(remote$url, path)
+ expect_equal(readLines(path), "b")
+ expect_equal(readLines(paste0(path, ".etag")), unname(tools::md5sum(remote$path)))
+})
+
+test_that("cached_download: falls back to the cached file if the remote can't be checked", {
+ remote <- local_remote("a")
+ path <- file.path(tempfile(), "file.txt")
+ local_mocked_bindings(remote_etag = function(url) unname(tools::md5sum(remote$path)))
+ tongfen:::cached_download(remote$url, path)
+
+ rm(list = remote$url, envir = tongfen:::tongfen_session)
+ writeLines("b", remote$path)
+ local_mocked_bindings(remote_etag = function(url) NULL)
+ expect_message(tongfen:::cached_download(remote$url, path), "using cached version")
+ expect_equal(readLines(path), "a")
+})
+
+test_that("cached_download: rejects downloads not matching the md5 ETag", {
+ remote <- local_remote("a")
+ path <- file.path(tempfile(), "file.txt")
+ local_mocked_bindings(remote_etag = function(url) strrep("0", 32))
+ expect_error(tongfen:::cached_download(remote$url, path), "corrupted")
+ expect_false(file.exists(path))
+})
+
+test_that("cached_download: check_remote = FALSE uses the cached file without contacting the remote", {
+ remote <- local_remote("a")
+ path <- file.path(tempfile(), "file.txt")
+ local_mocked_bindings(remote_etag = function(url) stop("remote should not be checked"))
+ expect_equal(tongfen:::cached_download(remote$url, path, check_remote = FALSE), path)
+ expect_equal(readLines(path), "a")
+ expect_false(file.exists(paste0(path, ".etag")))
+ # only the downloaded file ends up in the cache directory
+ expect_equal(list.files(dirname(path)), "file.txt")
+
+ # the cached file is used as is, also in a new session and if the remote changed
+ writeLines("b", remote$path)
+ tongfen:::cached_download(remote$url, path, check_remote = FALSE)
+ expect_equal(readLines(path), "a")
+
+ tongfen:::cached_download(remote$url, path, refresh = TRUE, check_remote = FALSE)
+ expect_equal(readLines(path), "b")
+})
+
+test_that("cached_download: failed downloads don't leave files in the cache", {
+ missing <- paste0("file://", normalizePath(tempdir(), winslash = "/"), "/does-not-exist.txt")
+ path <- file.path(tempfile(), "file.txt")
+ expect_error(suppressWarnings(tongfen:::cached_download(missing, path, check_remote = FALSE)))
+ expect_false(file.exists(path))
+ expect_equal(list.files(dirname(path)), character(0))
+})
+
+test_that("us_cache_dir: uses the given path and falls back to the tongfen cache directory", {
+ local_mocked_bindings(tongfen_cache_dir = function() "tongfen/cache")
+ expect_equal(tongfen:::us_cache_dir("some/path"), file.path("some/path", "us_data"))
+ expect_equal(tongfen:::us_cache_dir(NULL), file.path("tongfen/cache", "us_data"))
+ expect_equal(tongfen:::us_cache_dir(""), file.path("tongfen/cache", "us_data"))
+})
diff --git a/tests/testthat/test-correspondence-estimate.R b/tests/testthat/test-correspondence-estimate.R
index f1091d8..b237748 100644
--- a/tests/testthat/test-correspondence-estimate.R
+++ b/tests/testthat/test-correspondence-estimate.R
@@ -72,3 +72,38 @@ test_that("tongfen_tag_largest_overlap: tags each source region by its containin
expect_equal(tagged$t, c("t1", "t1", "t2", "t2"))
expect_equal(as.numeric(tagged[["...overlap_fraction"]]), rep(1, 4), tolerance = 1e-9)
})
+
+test_that("estimate_tongfen_correspondence: geometry column can go by any name", {
+ geo_a <- grid_4("idA")
+ geo_b <- st_sf(idB = c("b1", "b2"),
+ geometry = st_sfc(sq(0, 0, 2, 1), sq(0, 1, 2, 1), crs = 3347))
+ expected <- estimate_tongfen_correspondence(list(geo_a, geo_b), c("idA", "idB"),
+ tolerance = 0.01)
+
+ st_geometry(geo_a) <- "geom"
+ st_geometry(geo_b) <- "shape"
+ correspondence <- estimate_tongfen_correspondence(list(geo_a, geo_b), c("idA", "idB"),
+ tolerance = 0.01)
+ expect_equal(correspondence, expected)
+})
+
+test_that("estimate_tongfen_single_correspondence: robust option repairs invalid geometries", {
+ geo_a <- grid_4("idA")
+ geo_b <- st_sf(idB = c("b1", "b2"),
+ geometry = st_sfc(sq(0, 0, 2, 1), sq(0, 1, 2, 1), crs = 3347))
+ expected <- tongfen:::estimate_tongfen_single_correspondence(geo_a, geo_b, "idA", "idB",
+ tolerance = 0.01)
+
+ robust <- tongfen:::estimate_tongfen_single_correspondence(geo_a, geo_b, "idA", "idB",
+ tolerance = 0.01, robust = TRUE)
+ expect_equal(robust, expected)
+
+ # same region as b1, but the ring crosses over itself along the bottom edge
+ bowtie <- st_polygon(list(cbind(c(0, 2, 2, 0, 0, 1, 0), c(0, 0, 1, 1, 0, 0, 0))))
+ geo_invalid <- st_sf(idB = c("b1", "b2"),
+ geometry = st_sfc(bowtie, sq(0, 1, 2, 1), crs = 3347))
+ expect_false(all(st_is_valid(geo_invalid)))
+ robust <- tongfen:::estimate_tongfen_single_correspondence(geo_a, geo_invalid, "idA", "idB",
+ tolerance = 0.01, robust = TRUE)
+ expect_equal(robust, expected)
+})
diff --git a/tests/testthat/test-estimate.R b/tests/testthat/test-estimate.R
index c50efe7..48d75a3 100644
--- a/tests/testthat/test-estimate.R
+++ b/tests/testthat/test-estimate.R
@@ -45,3 +45,121 @@ test_that("tongfen_estimate preserves missing-value aggregation semantics", {
expect_true(is.na(keep_na$total))
expect_equal(drop_na$total, 5)
})
+
+test_that("tongfen_estimate ignores source regions that only touch the target", {
+ source <- sf::st_sf(
+ total = c(10, NA_real_),
+ geometry = sf::st_sfc(square(0, 1), square(1, 2), crs = 3347)
+ )
+ target <- sf::st_sf(
+ geometry = sf::st_sfc(square(0, 1), crs = 3347)
+ )
+ meta <- meta_for_additive_variables("synthetic", c(total = "total"))
+
+ # the second source region shares a boundary with the target but does not overlap it
+ expect_equal(tongfen_estimate(target, source, meta, na.rm = FALSE)$total, 10)
+ expect_equal(tongfen_estimate(target, source, meta, na.rm = TRUE)$total, 10)
+})
+
+test_that("tongfen_estimate handles intersections that are geometry collections", {
+ poly <- function(...) {
+ coords <- matrix(c(...), ncol = 2, byrow = TRUE)
+ sf::st_polygon(list(rbind(coords, coords[1, ])))
+ }
+ # comb shaped source, overlaps the target in two separate squares of area 0.5
+ # and touches it along a line in between
+ comb <- poly(0,1, 0,0.5, 1,0.5, 1,1, 2,1, 2,0.5, 3,0.5, 3,1, 3,2, 0,2)
+ source <- sf::st_sf(total = 100, geometry = sf::st_sfc(comb, crs = 3347))
+ target <- sf::st_sf(geometry = sf::st_sfc(poly(0,0, 3,0, 3,1, 0,1), crs = 3347))
+ meta <- meta_for_additive_variables("synthetic", c(total = "total"))
+
+ intersection <- sf::st_intersection(sf::st_geometry(source), sf::st_geometry(target))
+ expect_equal(as.character(sf::st_geometry_type(intersection)), "GEOMETRYCOLLECTION")
+
+ result <- tongfen_estimate(target, source, meta)
+
+ expect_equal(nrow(result), 1L)
+ expect_equal(result$total, 100 * 1 / as.numeric(sf::st_area(source)))
+})
+
+test_that("tongfen_estimate returns NA for target regions without overlap", {
+ source <- sf::st_sf(
+ total = 10,
+ geometry = sf::st_sfc(square(0, 1), crs = 3347)
+ )
+ target <- sf::st_sf(
+ region = c("inside", "outside"),
+ geometry = sf::st_sfc(square(0, 0.5), square(5, 6), crs = 3347)
+ )
+ meta <- meta_for_additive_variables("synthetic", c(total = "total"))
+
+ result <- tongfen_estimate(target, source, meta)
+ expect_equal(result$region, c("inside", "outside"))
+ expect_equal(result$total, c(5, NA_real_))
+
+ # no overlap at all
+ result <- tongfen_estimate(target[2, ], source, meta)
+ expect_s3_class(result, "sf")
+ expect_equal(nrow(result), 1L)
+ expect_true(is.na(result$total))
+})
+
+test_that("tongfen_estimate estimates parent weighted averages", {
+ source <- sf::st_sf(
+ hh = c(10, 30),
+ avg = c(2, 4),
+ geometry = sf::st_sfc(square(0, 1), square(1, 2), crs = 3347)
+ )
+ target <- sf::st_sf(
+ geometry = sf::st_sfc(square(0, 2), crs = 3347)
+ )
+ meta <- tibble::tibble(
+ variable = c("hh", "avg"), dataset = "synthetic", label = c("hh", "avg"),
+ type = "Manual", aggregation = c("Additive", "Average of hh"),
+ rule = c("Additive", "Average"), geo_dataset = "synthetic",
+ parent = c(NA, "hh")
+ )
+
+ result <- tongfen_estimate(target, source, meta)
+ expect_equal(result$hh, 40)
+ expect_equal(result$avg, 3.5)
+})
+
+test_that("tongfen_estimate complains about target columns that clash with the estimates", {
+ source <- sf::st_sf(
+ total = 10,
+ geometry = sf::st_sfc(square(0, 1), crs = 3347)
+ )
+ target <- sf::st_sf(
+ total = 1,
+ geometry = sf::st_sfc(square(0, 0.5), crs = 3347)
+ )
+ meta <- meta_for_additive_variables("synthetic", c(total = "total"))
+
+ expect_error(tongfen_estimate(target, source, meta), "already has columns named total")
+})
+
+test_that("tongfen_estimate takes averages over the regions that have a value with na.rm = TRUE", {
+ source <- sf::st_sf(
+ hh = c(10, 30, 60),
+ avg = c(2, NA, 4),
+ geometry = sf::st_sfc(square(0, 1), square(1, 2), square(2, 3), crs = 3347)
+ )
+ target <- sf::st_sf(
+ geometry = sf::st_sfc(square(0, 3), square(0.5, 2), crs = 3347)
+ )
+ meta <- tibble::tibble(
+ variable = c("hh", "avg"), dataset = "synthetic", label = c("hh", "avg"),
+ type = "Manual", aggregation = c("Additive", "Average of hh"),
+ rule = c("Additive", "Average"), geo_dataset = "synthetic",
+ parent = c(NA, "hh")
+ )
+
+ result <- tongfen_estimate(target, source, meta, na.rm = TRUE)
+ expect_equal(names(result), c("hh", "avg", "geometry"))
+ expect_equal(result$hh, c(100, 35))
+ # the second target only gets a value from half of the first source region
+ expect_equal(result$avg, c((2 * 10 + 4 * 60) / 70, 2))
+
+ expect_true(all(is.na(tongfen_estimate(target, source, meta, na.rm = FALSE)$avg)))
+})
diff --git a/tests/testthat/test-helpers.R b/tests/testthat/test-helpers.R
index 2914fd9..60b2c78 100644
--- a/tests/testthat/test-helpers.R
+++ b/tests/testthat/test-helpers.R
@@ -192,6 +192,21 @@ test_that("aggregate_correspondences: uses every input correspondence exactly on
expect_equal(nrow(result), 2L)
})
+test_that("aggregate_correspondences: only joins correspondences sharing an identifier", {
+ # ordered by size the first two correspondences have no identifier in common
+ cl <- list(
+ tibble(A = c("a1", "a2"), B = c("b1", "b2"), TongfenMethod = "statcan"),
+ tibble(B = c("b1", "b2", "b3", "b4"), C = c("c1", "c2", "c3", "c4"), TongfenMethod = "statcan"),
+ tibble(C = c("c1", "c2", "c3"), D = c("d1", "d2", "d3"), TongfenMethod = "statcan")
+ )
+ rlang::local_options(lifecycle_verbosity = "error")
+ result <- tongfen:::aggregate_correspondences(cl)
+ expect_equal(sort(names(result)), c("A", "B", "C", "D", "TongfenMethod"))
+ expect_equal(nrow(result), 2L)
+
+ expect_error(tongfen:::aggregate_correspondences(cl[c(1, 3)]), "common geographic identifier")
+})
+
# ── summarize_geometry_by_group ───────────────────────────────────────────────
test_that("summarize_geometry_by_group: matches grouped st_union", {
diff --git a/tests/testthat/test-proportional-reaggregate.R b/tests/testthat/test-proportional-reaggregate.R
index fc6a2ff..3f4a8a2 100644
--- a/tests/testthat/test-proportional-reaggregate.R
+++ b/tests/testthat/test-proportional-reaggregate.R
@@ -312,3 +312,63 @@ test_that("proportional_reaggregate: each category uses its own base variable",
expect_equal(result$v1, c(50, 0), tolerance = 1e-9)
expect_equal(result$v2, c(0, 50), tolerance = 1e-9)
})
+
+# ── Existing child values are compared to the parent per category ─────────────
+
+test_that("proportional_reaggregate: several categories with existing child values match their parent totals", {
+ parent <- tibble(
+ parent_id = "P1",
+ pop = 100,
+ x = 50,
+ y = 6
+ ) %>% as_point_sf()
+
+ child <- tibble(
+ child_id = c("C1", "C2", "C3"),
+ parent_id = "P1",
+ pop = c(20, 30, 50),
+ x = c(10, 15, 20),
+ y = c(1, 2, 3)
+ ) %>% as_point_sf()
+
+ result <- proportional_reaggregate(
+ child, parent,
+ geo_match = c("parent_id" = "parent_id"),
+ categories = c("x", "y"),
+ base = "pop"
+ ) %>%
+ sf::st_drop_geometry() %>%
+ arrange(.data$child_id)
+
+ # x is 5 short of the parent and gets topped up by population share, y already matches
+ expect_equal(result$x, c(11, 16.5, 22.5), tolerance = 1e-9)
+ expect_equal(result$y, c(1, 2, 3), tolerance = 1e-9)
+})
+
+test_that("proportional_reaggregate: missing child values are filled so that children sum to parent", {
+ parent <- tibble(
+ parent_id = "P1",
+ pop = 300,
+ x = 50
+ ) %>% as_point_sf()
+
+ child <- tibble(
+ child_id = c("C1", "C2", "C3"),
+ parent_id = "P1",
+ pop = c(100, 100, 100),
+ x = c(10, NA, 20)
+ ) %>% as_point_sf()
+
+ result <- proportional_reaggregate(
+ child, parent,
+ geo_match = c("parent_id" = "parent_id"),
+ categories = "x",
+ base = "pop"
+ ) %>%
+ sf::st_drop_geometry() %>%
+ arrange(.data$child_id)
+
+ # the 20 missing from the children get distributed evenly
+ expect_equal(result$x, c(10, 0, 20) + 20/3, tolerance = 1e-9)
+ expect_equal(sum(result$x), 50, tolerance = 1e-9)
+})
diff --git a/vignettes/tongfen-ca-estimate.Rmd b/vignettes/tongfen-ca-estimate.Rmd
index ba3bcb9..bd73a86 100644
--- a/vignettes/tongfen-ca-estimate.Rmd
+++ b/vignettes/tongfen-ca-estimate.Rmd
@@ -24,7 +24,7 @@ library(ggplot2)
The `tongfen_estimate_ca_census` function implements the full pipeline to estimate census data on custom geographies.
-As an example, we estimate the share of people in low income in Vancouver's skytrain station neighbourhoods. The station neighbourhoods are available as part of the `cancenus` package.
+As an example, we estimate the share of people in low income in Vancouver's skytrain station neighbourhoods. The station neighbourhoods are available as part of the `cancensus` package.
```{r}
station_buffers <- cancensus::COV_SKYTRAIN_STATIONS
diff --git a/vignettes/tongfen.Rmd b/vignettes/tongfen.Rmd
index 27426c9..30846ba 100644
--- a/vignettes/tongfen.Rmd
+++ b/vignettes/tongfen.Rmd
@@ -48,7 +48,7 @@ data <- years %>%
}) %>% setNames(years)
```
-Plotting the cenus tracts for our four census years shows how census tracts changed over the years.
+Plotting the census tracts for our four census years shows how census tracts changed over the years.
```{r}
data %>%
@@ -60,9 +60,9 @@ data %>%
labs(title="Vancouver census tracts",caption="StatCan Census 2001-2016")
```
-For this example we will estimate the correspondence between these regions from the geographic data using the `estimate_tongfen_correspondence` function. Unfortunately this is not an exact science, for example over the years census regions get adjusted to better align with the road network. Other harmless boundary adjustemens can happen along water boundaries, or re-jigging boundaries in unpopulated areas.
+For this example we will estimate the correspondence between these regions from the geographic data using the `estimate_tongfen_correspondence` function. Unfortunately this is not an exact science, for example over the years census regions get adjusted to better align with the road network. Other harmless boundary adjustments can happen along water boundaries, or re-jigging boundaries in unpopulated areas.
-We are going to impose a tolerance of 200m, where we are calling two census tract the same if they differ by no more than 200m. We are specifying that these calculations should be carried out in the Statistics Canada Lambert (EPSG:3347) refernce system with units metres.
+We are going to impose a tolerance of 200m, where we are calling two census tract the same if they differ by no more than 200m. We are specifying that these calculations should be carried out in the Statistics Canada Lambert (EPSG:3347) reference system with units metres.
```{r}
correspondence <- estimate_tongfen_correspondence(data, geo_identifiers,
@@ -124,7 +124,7 @@ years %>%
```
## Population change
-It's time to go back to our original goal of mapping population change. For this we need to specify how to aggregate up the population data, which is by simply adding them up. The `meta_for_additive_variables` convenience function generates the appropriate metatdata that specifies how to deal with this data.
+It's time to go back to our original goal of mapping population change. For this we need to specify how to aggregate up the population data, which is by simply adding them up. The `meta_for_additive_variables` convenience function generates the appropriate metadata that specifies how to deal with this data.
```{r}
meta <- meta_for_additive_variables(years,"Population")
diff --git a/vignettes/tongfen_anomalies.Rmd b/vignettes/tongfen_anomalies.Rmd
new file mode 100644
index 0000000..fe2e89a
--- /dev/null
+++ b/vignettes/tongfen_anomalies.Rmd
@@ -0,0 +1,202 @@
+---
+title: "Geocoding anomalies in TongFen timelines"
+author: "Jens von Bergmann"
+date: "`r Sys.Date()`"
+output: rmarkdown::html_vignette
+vignette: >
+ %\VignetteIndexEntry{Geocoding anomalies in TongFen timelines}
+ %\VignetteEngine{knitr::rmarkdown}
+ %\VignetteEncoding{UTF-8}
+---
+
+```{r setup, include = FALSE}
+knitr::opts_chunk$set(
+ message = FALSE,
+ warning = FALSE,
+ collapse = TRUE,
+ eval = nzchar(Sys.getenv("COMPILE_VIG")),
+ comment = "#>"
+)
+```
+
+TongFen makes data on different geographies comparable by aggregating it up to a common geography. The result is only as good as the geocoding that assigned the underlying data to geographic regions in the first place. Geocoding changes over time, and the same dwelling units, and the people living in them, can get assigned to different neighbouring regions in different years. In a timeline on a common geography this shows up as a surprising drop in one region that is offset by a jump in a neighbouring region.
+
+The fix is in the spirit of TongFen, joining the affected regions gives a slightly coarser geography on which the data is consistent over time. The functions in this vignette look for such patterns and join the regions on demand. The method is explained in more detail in a [blog post](https://doodles.mountainmath.ca/posts/2024-07-26-geocoding-errors-in-aggregate-data/), this vignette follows the example from that post.
+
+```{r}
+library(dplyr)
+library(tidyr)
+library(ggplot2)
+library(cancensus)
+library(sf)
+library(tongfen)
+# cancensus::set_api_key("")
+```
+
+## Population timelines for Toronto
+
+As an example we take the population from the 1971 through 2011 censuses that Statistics Canada tabulated on 2016 dissemination areas, together with the 2016 population. All data comes on the same geography, so there is no need to TongFen, but the data for the earlier years is geocoded from the road network and block face of the time, which does not always match up with 2016 dissemination areas.
+
+```{r}
+years <- c(1971,seq(1981,2011,5))
+vectors <- c(setNames(paste0("v_CA",years,"x16_1"),years),"2016"="v_CA16_1")
+timeline <- names(vectors)
+
+toronto <- get_census("CA16CT",regions=list(CSD="3520005"),vectors=vectors,
+ level="DA",geo_format="sf",quiet=TRUE) %>%
+ select(GeoUID,all_of(timeline)) %>%
+ mutate(across(all_of(timeline),\(x) coalesce(x,0)))
+```
+
+Dissemination areas without population in a given year come back as missing values. Changes from or to a missing value are never considered surprising, so we set them to zero to mark these as areas where nobody got counted.
+
+The area around Crescent Town shows what the problem looks like.
+
+```{r fig.width=7, fig.height=3.5}
+crescent_town <- c("35204370","35204765")
+
+plot_timelines <- function(data) {
+ data %>%
+ st_drop_geometry() %>%
+ pivot_longer(all_of(timeline),names_to="Year",values_to="Population") %>%
+ ggplot(aes(x=Year,y=Population,colour=GeoUID,group=GeoUID)) +
+ geom_line() +
+ geom_point() +
+ scale_y_continuous(labels=scales::comma,limits=c(0,NA))
+}
+
+toronto %>%
+ filter(GeoUID %in% crescent_town) %>%
+ plot_timelines() +
+ labs(title="Population in two neighbouring dissemination areas")
+```
+
+The population jumps back and forth between the two areas, while the sum of the two is fairly steady from 1981 on. People did not move back and forth, their homes got geocoded to a different dissemination area in different years.
+
+## Detecting anomalies
+
+`tongfen_detect_anomalies` lists the regions with surprising drops. Only decreases are surprising, and a decrease needs to be large in both relative and absolute terms. For each of these candidate regions it finds the neighbouring region that takes away most of the surprise when both are joined, and checks if that reduction is large enough to justify joining them. Our regions are identified by their `GeoUID` instead of the `TongfenID` the function looks for by default.
+
+```{r}
+anomalies <- tongfen_detect_anomalies(toronto,timeline,id="GeoUID",total_surprise_cutoff=0.4)
+
+anomalies %>% filter(GeoUID %in% crescent_town)
+```
+
+Both areas are candidates, and each one is the neighbour that best explains the surprising drops of the other. Not all candidates find a neighbour to pair up with. Population does drop for real, for example when a site gets cleared for redevelopment, and such regions are left alone.
+
+```{r}
+anomalies %>% count(join)
+```
+
+## Joining regions
+
+`tongfen_anomaly_joins` joins the regions that qualify and looks again, joined regions can have surprising drops that are complemented by another neighbour. This repeats until there are no more regions left to join. The result lists the regions that got joined together with the identifier of the joined region they are now part of and the round in which they first got joined.
+
+```{r}
+joins <- tongfen_anomaly_joins(toronto,timeline,id="GeoUID",total_surprise_cutoff=0.4)
+
+joins %>% filter(GeoUID %in% crescent_town)
+```
+
+`tongfen_join_regions` applies the joins to the data, aggregating the variables and geometries of the regions that get joined and leaving all others as they are.
+
+```{r}
+toronto_joined <- tongfen_join_regions(toronto,joins,id="GeoUID")
+
+c(original=nrow(toronto),joined=nrow(toronto_joined))
+```
+
+```{r fig.width=7, fig.height=3.5}
+toronto_joined %>%
+ filter(GeoUID %in% crescent_town) %>%
+ plot_timelines() +
+ labs(title="Population in the joined region")
+```
+
+The map shows the regions that got joined around Crescent Town.
+
+```{r fig.width=7, fig.height=5}
+bbox <- toronto %>% filter(GeoUID %in% crescent_town) %>% st_buffer(1500) %>% st_bbox()
+
+ggplot(toronto_joined %>% mutate(joined=GeoUID %in% joins$GeoUID_joined)) +
+ geom_sf(aes(fill=joined),linewidth=0.1) +
+ geom_sf(data=toronto,fill=NA,linewidth=0.1,linetype="dotted") +
+ scale_fill_manual(values=c("TRUE"="steelblue","FALSE"="whitesmoke"),guide="none") +
+ coord_sf(datum=NA,xlim=bbox[c("xmin","xmax")],ylim=bbox[c("ymin","ymax")]) +
+ labs(title="Joined regions around Crescent Town",
+ caption="Joined regions in blue, original dissemination areas dotted")
+```
+
+## Tuning
+
+Joining regions trades geographic detail for consistency over time, and how to best make that trade depends on the data and the application. The parameters are documented in `tongfen_detect_anomalies`, the most important ones are
+
+* `rel_scale` and `abs_scale`, the relative and absolute decrease at which a change is half way to being fully surprising. The defaults of a 25% drop and a drop of 200 are tuned to population counts in regions of the size of dissemination areas.
+* `total_surprise_cutoff`, how surprising the timeline of a region needs to be to become a candidate. The default of 0.75 is conservative, above we used 0.4 to also pick up less pronounced cases.
+* `cutoff_fact`, `surprise_reduction_const` and `sum_fact` determine how much of the surprise a neighbour needs to take away for the regions to get joined.
+
+```{r}
+c(0.4,0.6,0.75) %>%
+ lapply(\(cutoff) tibble(total_surprise_cutoff=cutoff,
+ regions_joined=tongfen_anomaly_joins(toronto,timeline,id="GeoUID",
+ total_surprise_cutoff=cutoff) %>%
+ nrow())) %>%
+ bind_rows()
+```
+
+Neighbours are by default determined by intersecting the geometries of the regions. This can miss neighbours if the geometries have been simplified, in that case the `neighbours` argument takes a table with the identifiers of neighbouring regions or a neighbours list from the **spdep** package.
+
+## Anomalies in TongFen data
+
+The functions work the same way on data on a common geography built by TongFen, where the regions are identified by their `TongfenID`. As an example we look at the dissemination area level population in the City of Vancouver for the 2001 through 2021 censuses.
+
+```{r}
+regions <- list(CSD="5915022")
+datasets <- c("CA01","CA06","CA11","CA16","CA21")
+meta <- meta_for_additive_variables(datasets,"Population")
+
+vancouver <- get_tongfen_ca_census(regions=regions,meta=meta,level="DA",base_geo="CA21",quiet=TRUE)
+
+joins <- tongfen_anomaly_joins(vancouver,paste0("Population_",datasets),total_surprise_cutoff=0.4)
+joins
+```
+
+Passing the metadata to `tongfen_join_regions` makes sure the variables get aggregated the right way, numeric variables that are not part of the metadata are assumed to be additive.
+
+```{r}
+vancouver_joined <- tongfen_join_regions(vancouver,joins,meta)
+```
+
+The `TongfenUID` of the joined regions lists all the dissemination areas they are made up of.
+
+```{r}
+vancouver_joined %>%
+ st_drop_geometry() %>%
+ filter(TongfenID %in% joins$TongfenID_joined) %>%
+ select(TongfenID,TongfenUID,starts_with("Population"))
+```
+
+Variables that are not additive, like averages, can only be aggregated this way if the variable they are averaged over is part of the data. The alternative that always works is to join the regions in the correspondence the common geography was built from, and use the joined correspondence to aggregate the original data. This also is the way to use the joins for data other than the one that was used to detect the anomalies, for example to get average rents on the corrected geography.
+
+```{r}
+correspondence <- get_tongfen_correspondence_ca_census(geo_datasets=datasets,regions=regions,
+ level="DA",quiet=TRUE) %>%
+ tongfen_join_correspondence(joins)
+
+rent_meta <- meta_for_ca_census_vectors(c(rent_2006="v_CA06_2050",rent_2016="v_CA16_4901"))
+
+rent_data <- c("CA06","CA16") %>%
+ lapply(\(ds) get_census(ds,regions=regions,level="DA",labels="short",quiet=TRUE,
+ vectors=rent_meta %>% filter(geo_dataset==ds) %>% pull(variable),
+ geo_format=if (ds=="CA16") "sf" else NA) %>%
+ rename(!!paste0("GeoUID",ds):="GeoUID")) %>%
+ setNames(c("CA06","CA16"))
+
+rents <- tongfen_aggregate(rent_data,correspondence,rent_meta,base_geo="CA16")
+
+rents %>%
+ st_drop_geometry() %>%
+ filter(TongfenID %in% joins$TongfenID_joined) %>%
+ select(TongfenID,rent_2006,rent_2016)
+```