Repository navigation
v0.3.9 - #20
Merged
Merged
v0.3.9#20
Conversation
aggregate_correspondences joined the per-step correspondences in order of size, so with three or more datasets consecutive tables could have no identifier in common and got cross joined. That triggered a dplyr deprecation warning and large intermediate tables. Tables now only get joined once they share an identifier with the ones joined so far. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Statistics Canada put the correspondence files behind a browser check that blocks programmatic downloads, breaking method = "statcan". The files are now hosted as parquet on S3, built by data-raw/statcan_correspondence.R from manually downloaded originals. Cached files are checked against the mirror via ETag once per session and downloaded again if they changed. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The master branch has been removed, main is the default branch. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…dit fixes Wrong results - proportional_reaggregate compared existing child values to the parent total across all categories instead of per category, and a missing child value discarded the values of its siblings - tongfen_estimate underestimated values when an intersection came back as a geometry collection, and picked up NA from source regions that only touch the target - averages were scaled by their parent more than once when the metadata lists the same variable for several datasets, tongfen_aggregate now only uses the metadata of the dataset being aggregated - averages aggregated with na.rm = TRUE kept the parent of regions with a missing value in the denominator - "Average to" variables sharing a parent overwrote each other's base Errors - estimate_tongfen_correspondence required the geometry column to be named geometry, and the robust option failed on more than one region - tongfen_estimate errored out when no target region overlaps the source, and gave an obscure error on clashing column names Behaviour - refresh is passed through to the StatCan correspondence download - tongfen_estimate_ca_census returns visibly - US relationship files are downloaded to a temporary file first and cached via the shared tongfen cache directory - fix function names in the get_tongfen_ca_census_ct_from_da deprecation Docs and housekeeping - missing \value sections, broken or misleading examples, README and get_tongfen_us_census docs - remove dead code in tongfen_ca.R - require dplyr >= 1.1.0, which the package already relied on - NEWS and cran-comments for 0.3.9 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Geocoding varies between censuses, the same dwellings can get assigned to different neighbouring regions in different years. On a common geography this shows up as a surprising drop in one region that is offset by a jump in a neighbour. Adds experimental functions to check for this and to join the affected regions on demand, following the method from https://doodles.mountainmath.ca/posts/2024-07-26-geocoding-errors-in-aggregate-data/ - tongfen_detect_anomalies lists candidate regions with their best neighbour - tongfen_anomaly_joins iterates joins until no more regions qualify - tongfen_join_regions applies the joins to data on a common geography - tongfen_join_correspondence applies them to a correspondence for use with tongfen_aggregate The neighbour graph is contracted between rounds instead of recomputed from the joined geometries, so joined regions keep all neighbours of their parts. Missing values are never surprising and don't make up for a drop in a neighbour, and ties are broken by identifier so results don't depend on the order of the regions. Hoists intersects_pairs into helpers.R for reuse, adds a vignette and tests including a comparison against the original implementation. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01A4ixWqz6d5wpYg2DkoT4pm
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Release branch for v0.3.9.
Fixes #19.
Closes #21.
Major changes
tongfen_detect_anomalieslists the candidate regions with their best neighbour.tongfen_anomaly_joinsdetermines the regions to join, iterating until no more regions qualify.tongfen_join_regionsjoins them in data that already is on a common geography.tongfen_join_correspondencejoins them in a correspondence for use withtongfen_aggregate.method = "statcan"(StatCan disabled programmatic download of correspondence files #19).tongfen.statcan_correspondence_urloption.statcan_correspondence_*.csvfiles in the tongfen cache directory are no longer used and can be removed.Fixes with changed results
proportional_reaggregategave wrong results when the finer level data already has values for the categories to reaggregate. The values were compared to the parent total across all categories instead of per category, and a missing value in one child region discarded the existing values of all its siblings.tongfen_estimateunderestimated values when the intersection of a source and a target region is a geometry collection.tongfen_aggregateandaggregate_data_with_metascaled averages by their parent variable more than once when the metadata lists the same variable name for several datasets, as is common with US census data.tongfen_aggregatenow only uses the metadata of the dataset being aggregated, conflicting aggregation rules for the same variable are an error.na.rm = TRUEare now taken over the regions that have a value. The parent variable of regions with a missing average still counted toward the total, pulling the result toward zero.base_<variable>instead of after the parent.Minor changes
tongfen_estimatewithna.rm = FALSEno longer returnsNAfor target regions that only share a boundary with a source region with missing values, returnsNAfor target regions that don't overlap the source instead of erroring out when none of them do, and gives a clear error whentargetalready has a column named like one of the variables to estimate.estimate_tongfen_correspondenceno longer requires the geometry column to be namedgeometry.refresh = TRUEinget_tongfen_correspondence_ca_censusandget_tongfen_ca_censusnow also refreshes the cached StatCan correspondence files.tongfen.cache_pathenvironment variable and thecustom_data_pathoption.tongfen_estimate_ca_censusreturns its result visibly.\valuesections, and thetongfen_aggregateexample now passes a named list of datasets matching the metadata.masterbranch.Notes
nanoparquet(andtools) in Imports, dplyr requirement raised to 1.1.0, which the package already relied on.data-raw/statcan_correspondence.Rbuilds the parquet files from the manually downloaded StatCan originals. The eight files are hosted underhttps://mountainmath.s3.ca-central-1.amazonaws.com/tongfen/statcan_correspondence/v1/.cran-comments.mdis updated for the CRAN submission.Testing
R CMD check --as-cranon macOS with R 4.6.0, with all six vignettes rebuilt against live data: 0 errors, 0 warnings, 0 notes. This ran before the spelling and pkgdown commits, which only touch documentation.get_tongfen_correspondence_ca_censuswithmethod = "statcan"for CT and DA levels across the 2001 to 2021 censuses.🤖 Generated with Claude Code
https://claude.ai/code/session_01A4ixWqz6d5wpYg2DkoT4pm