Skip to content

drop --spiked - #17

Merged
harmsm merged 6 commits into
harmslab:mainfrom
harmsm:main
Sep 23, 2026
Merged

harmsm merged 6 commits into
harmslab:mainfrom
harmsm:main

Conversation

@harmsm

@harmsm harmsm commented Sep 23, 2026

Copy link
Copy Markdown
Contributor
  • Big change is dropping --spiked from tfs-configure-fit. This now takes a yaml file describing the library as input. This is in preparation for a larger model change that will change how spiked and congressional transformation are treated. For now, the logic is the same.
  • Updated process_raw to take sample_ln_cfu
  • Fixed some test failures and installation hangups

harmsm and others added 6 commits September 22, 2026 08:03
counts_to_lncfu required linear-space sample_cfu/sample_cfu_std, even
though the calculation is naturally done in log space and downstream
consumers work with ln_cfu.

counts_to_lncfu now reads sample_ln_cfu/sample_ln_cfu_std, inferring
them via get_scaled_cfu from sample_ln_cfu_var or from sample_cfu plus
sample_cfu_std/sample_cfu_var when absent (log-space columns win).
Genotype ln_cfu/ln_cfu_var are computed directly in log space; cfu and
cfu_var are derived. New helper get_sample_ln_cfu is called early by
tfs-process-counts and tfs-process-presplit so bad sample sheets fail
before any counts are read. get_scaled_cfu gains a `prefix` argument.

Existing sample sheets with sample_cfu/sample_cfu_std keep working and
give identical results; outputs gain sample_ln_cfu/sample_ln_cfu_std.

Verified: new tests for inference paths, precedence, prefix handling
and log/linear agreement; tests/tfscreen/process_raw and util pass.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
numpyro >= 0.20 validates distribution arguments by default and treats
HalfNormal's support as open, so CI (numpyro 0.22 / jax 0.11) failed
seven tests that pass locally on numpyro 0.19.

- horseshoe: floor the distance-dependent lambda scale at 1e-12;
  exp(-d/d0) underflowed to 0 for distant pairs, an invalid HalfCauchy
  scale.
- test_prediction: build fake one-draw posteriors with each site's
  support.feasible_like instead of zeros, so positive-support sites
  (noise sigmas) get valid values.
- test_loglikelihood: check the HalfNormal density at x > 0 instead of
  at the now-excluded boundary x = 0.

Verified: full tests/tfscreen passes under numpyro 0.22.0/jax 0.11.2
(4414 passed) and the touched files pass under numpyro 0.19.0.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
`--spiked` did two jobs at once -- it set both congression purity and the
ln_cfu0 prior class -- and it could not express the fact that the spiked
controls are also encoded in the bulk sub-libraries. wt and each spiked
single mutant arrive both as a monoclonal spike and as ordinary library
members, and sequence cannot tell the two apart, so a hand-typed list of
"spiked" genotypes is wrong at the level of what it can represent. In
examples/simulate-and-analyze the hand-written list had already drifted from
the config it described: it omitted M42I/H74A/K84L, which is in spiked_seqs.

tfs-configure-model now takes --library_config: the library YAML that
describes the screened library, the same file handed to tfs-process-fastq.
Required whenever growth_df is given; refused for a binding-only model.

- genetics/library_design.py gains library_composition_table() plus
  read/write_library_composition(). The table carries is_wt,
  in_spiked_origin, pool_fraction and bulk_fraction per genotype. The module
  also brings expected_library_composition(), scale_library_design() and
  estimate_library_mixture() -- the last estimates a realized library_mixture
  from an observed abundance table, which is the instrument for checking a
  library's composition against its design.
- configure-model writes {out_prefix}_library.csv and records it, with the
  declared library_mixture, in the config. tfs-fit-model reads the snapshot,
  never the YAML -- the same pattern as the priors and guesses CSVs, and the
  place to override a design number.
- ModelOrchestrator takes library_file and derives today's spiked_genotypes
  from in_spiked_origin, so the model backend is unchanged. A library-derived
  spike with no growth data is reported, not fatal (spikes drop out of real
  data); the hand-supplied list still raises. The two are mutually exclusive
  and configs carrying spiked_genotypes still load.
- configure-model fails if any genotype in growth_df/presplit_df/
  base_growth_df is absent from the library. This catches a residue-numbering
  or wt-sequence mismatch between the library description and the reads it was
  called against -- the presplit failure mode -- on the first run.
- tfs-fit-genotypes and tfs-build-empirical move from --spiked_file to
  --library_config. tfs-setup-grid resolves library_config like its other
  path arguments, so grid runs can pass it per combination.

pool_fraction/bulk_fraction are snapshotted but not yet consumed: bulk_fraction
is the f_g the observable-level congression mixture will use. library_mixture
is declared-only -- no upstream artifact ever saw it, so nothing can check it,
and recording it verbatim is the whole defense. Documented that one
run_config.yaml must be shared by process-fastq, simulate and configure-model.

Also here: CHANGELOG.md, carrying the reconstructed release history and an
[Unreleased] section covering this change and the earlier log-space
counts_to_lncfu work; and a fix to LibraryManager._get_spiked_seqs's
docstring, which claimed the function expands a spiked double into its
singles and wt. It does not, and library composition depends on that.

Verification: full unit suite (4463 passed, 6 skipped) and smoke tests with
--runslow (37 passed). New tests cover the composition table, the snapshot
round trip, mask equivalence against the legacy --spiked list, the
absent-spike and numbering-shift cases, and legacy config loading.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@harmsm
harmsm merged commit 0ae9179 into harmslab:main Sep 23, 2026
6 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant