Repository navigation
drop --spiked - #17
Merged
Merged
Conversation
counts_to_lncfu required linear-space sample_cfu/sample_cfu_std, even though the calculation is naturally done in log space and downstream consumers work with ln_cfu. counts_to_lncfu now reads sample_ln_cfu/sample_ln_cfu_std, inferring them via get_scaled_cfu from sample_ln_cfu_var or from sample_cfu plus sample_cfu_std/sample_cfu_var when absent (log-space columns win). Genotype ln_cfu/ln_cfu_var are computed directly in log space; cfu and cfu_var are derived. New helper get_sample_ln_cfu is called early by tfs-process-counts and tfs-process-presplit so bad sample sheets fail before any counts are read. get_scaled_cfu gains a `prefix` argument. Existing sample sheets with sample_cfu/sample_cfu_std keep working and give identical results; outputs gain sample_ln_cfu/sample_ln_cfu_std. Verified: new tests for inference paths, precedence, prefix handling and log/linear agreement; tests/tfscreen/process_raw and util pass. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
numpyro >= 0.20 validates distribution arguments by default and treats HalfNormal's support as open, so CI (numpyro 0.22 / jax 0.11) failed seven tests that pass locally on numpyro 0.19. - horseshoe: floor the distance-dependent lambda scale at 1e-12; exp(-d/d0) underflowed to 0 for distant pairs, an invalid HalfCauchy scale. - test_prediction: build fake one-draw posteriors with each site's support.feasible_like instead of zeros, so positive-support sites (noise sigmas) get valid values. - test_loglikelihood: check the HalfNormal density at x > 0 instead of at the now-excluded boundary x = 0. Verified: full tests/tfscreen passes under numpyro 0.22.0/jax 0.11.2 (4414 passed) and the touched files pass under numpyro 0.19.0. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
`--spiked` did two jobs at once -- it set both congression purity and the
ln_cfu0 prior class -- and it could not express the fact that the spiked
controls are also encoded in the bulk sub-libraries. wt and each spiked
single mutant arrive both as a monoclonal spike and as ordinary library
members, and sequence cannot tell the two apart, so a hand-typed list of
"spiked" genotypes is wrong at the level of what it can represent. In
examples/simulate-and-analyze the hand-written list had already drifted from
the config it described: it omitted M42I/H74A/K84L, which is in spiked_seqs.
tfs-configure-model now takes --library_config: the library YAML that
describes the screened library, the same file handed to tfs-process-fastq.
Required whenever growth_df is given; refused for a binding-only model.
- genetics/library_design.py gains library_composition_table() plus
read/write_library_composition(). The table carries is_wt,
in_spiked_origin, pool_fraction and bulk_fraction per genotype. The module
also brings expected_library_composition(), scale_library_design() and
estimate_library_mixture() -- the last estimates a realized library_mixture
from an observed abundance table, which is the instrument for checking a
library's composition against its design.
- configure-model writes {out_prefix}_library.csv and records it, with the
declared library_mixture, in the config. tfs-fit-model reads the snapshot,
never the YAML -- the same pattern as the priors and guesses CSVs, and the
place to override a design number.
- ModelOrchestrator takes library_file and derives today's spiked_genotypes
from in_spiked_origin, so the model backend is unchanged. A library-derived
spike with no growth data is reported, not fatal (spikes drop out of real
data); the hand-supplied list still raises. The two are mutually exclusive
and configs carrying spiked_genotypes still load.
- configure-model fails if any genotype in growth_df/presplit_df/
base_growth_df is absent from the library. This catches a residue-numbering
or wt-sequence mismatch between the library description and the reads it was
called against -- the presplit failure mode -- on the first run.
- tfs-fit-genotypes and tfs-build-empirical move from --spiked_file to
--library_config. tfs-setup-grid resolves library_config like its other
path arguments, so grid runs can pass it per combination.
pool_fraction/bulk_fraction are snapshotted but not yet consumed: bulk_fraction
is the f_g the observable-level congression mixture will use. library_mixture
is declared-only -- no upstream artifact ever saw it, so nothing can check it,
and recording it verbatim is the whole defense. Documented that one
run_config.yaml must be shared by process-fastq, simulate and configure-model.
Also here: CHANGELOG.md, carrying the reconstructed release history and an
[Unreleased] section covering this change and the earlier log-space
counts_to_lncfu work; and a fix to LibraryManager._get_spiked_seqs's
docstring, which claimed the function expands a spiked double into its
singles and wt. It does not, and library composition depends on that.
Verification: full unit suite (4463 passed, 6 skipped) and smoke tests with
--runslow (37 passed). New tests cover the composition table, the snapshot
round trip, mask equivalence against the legacy --spiked list, the
absent-spike and numbering-shift cases, and legacy config loading.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
--spikedfromtfs-configure-fit. This now takes a yaml file describing the library as input. This is in preparation for a larger model change that will change how spiked and congressional transformation are treated. For now, the logic is the same.