Skip to content

Monthly Conversion: Antiobiotic Data #2 - #94

Merged
KateJohnson merged 2 commits into
mainfrom
data-generation-monthly-antibiotics
Sep 5, 2026
Merged

KateJohnson merged 2 commits into
mainfrom
data-generation-monthly-antibiotics

Conversation

@ahill187

@ahill187 ahill187 commented Aug 12, 2026 •

Copy link
Copy Markdown
Collaborator

Description

This is a continuation of #82: Monthly Conversion: Antibiotic Exposure. PR #82 was closed by mistake, without update the *.csv files. I reverted that PR in PR #93.

This PR is thus everything from PR #82 , plus midtrends.csv has been renamed to antibiotic_predictions.csv and the the extra files deleted.

Type of Change

  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to not work as expected)
  • Documentation (updated docstrings or README files)
  • Tests (added new tests or modified current test suite)
  • Refactoring (changes in the design of the code that don't change the functionality)

Tests

Please describe the tests that you ran to verify your changes. Provide instructions so we can reproduce. Please also list any relevant details for your test configuration.

  • Test A
  • Test B

Linked PRs / Issues

Original PR: #82
Revert: #93

Checklist:

  • I have performed a self-review of my code
  • I have documented my code with appropriate docstrings
  • I have made corresponding changes to the documentation / README files
  • I have added tests that prove my fix is effective or that my feature works
  • New and existing unit tests pass locally with my changes

Copilot AI lite review requested due to automatic review settings August 12, 2026 20:06

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR updates the antibiotic exposure model to use the newer antibiotic predictions dataset (grouped by timepoint + sex) and aligns the public API/tests with the renamed data accessor.

Changes:

  • Renames the internal exposure dataset handle from mid_trends to data and updates copying logic accordingly.
  • Switches the data loader from processed_data/.../midtrends.csv to processed_data/.../antibiotic_predictions.csv (with the new n_abx_μ column).
  • Removes the old midtrends.csv files and updates the constructor test to reference .data.

Reviewed changes

Copilot reviewed 4 out of 4 changed files in this pull request and generated 1 comment.

File Description
tests/test_antibiotic_exposure.py Updates constructor test to use the renamed .data grouping accessor.
leap/processed_data/time_delta_365/midtrends.csv Deletes the legacy midtrends dataset.
leap/processed_data/time_delta_30/midtrends.csv Deletes the legacy midtrends dataset.
leap/antibiotic_exposure.py Renames API surface from mid_trends to data and changes loader to antibiotic_predictions.csv / n_abx_μ.
Suppressed comments (2)

leap/antibiotic_exposure.py:121

  • antibiotic_predictions.csv has sex values like F/M (see processed_data/time_delta_*/antibiotic_predictions.csv), but the API and tests treat sex as 0/1. As loaded, df.groupby(["timepoint", "sex"]) will be keyed by strings, so get_group((..., 0)) and downstream code using int(sex) won't match. Map the CSV sex values to 0/1 (and validate unexpected values) before grouping.
        df = pd.read_csv(
            get_data_path(f"processed_data/{time_delta_tag}/antibiotic_predictions.csv"),
            parse_dates=["timepoint"]
        )
        grouped_df = df.groupby(["timepoint", "sex"])

leap/antibiotic_exposure.py:160

  • In the fixyear non-numeric branch, self.data is grouped by timepoint (a parsed datetime), but this uses birth_year (an int) as the group key. This will raise a KeyError at runtime when fixyear is set and not numeric. Convert birth_year to the same dt.datetime(birth_year, 1, 1) key used elsewhere.
                μ = max(
                    self.data.get_group((birth_year, int(sex)))["n_abx_μ"].iloc[0],
                    self.parameters["βfloor"]
                )
                p = self.parameters["θ"] / (self.parameters["θ"] + μ)

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Each entry is a dataframe with a single row with the following columns:

* ``year (int)``: The calendar year, e.g. ``2024``.
* ``timepoint (dt.datetime)``: The date and time, e.g. ``2024``.
@KateJohnson
KateJohnson self-requested a review September 1, 2026 20:10
antibiotic_predictions.csv encodes sex as F/M strings, but code expected
0/1 ints from the old midtrends.csv, causing get_group() KeyErrors.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
@KateJohnson
KateJohnson merged commit 9b46122 into main Sep 5, 2026
1 check passed
@KateJohnson
KateJohnson deleted the data-generation-monthly-antibiotics branch September 5, 2026 04:03
@ahill187
ahill187 restored the data-generation-monthly-antibiotics branch September 8, 2026 22:43
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants