From CROWN ntuples to histograms, background estimates, combine datacards, fits, plots,
ML training folds and measurements (correctionlib payloads). Columnar (uproot, pandas,
numexpr, numpy), no ROOT needed except for shapesmith fit, which calls combine in a CMSSW
subshell.
An analysis is a Python package that builds a shapesmith.model.Analysis from a small run
YAML; ShapeSmith runs the steps. Example analysis:
BBTauTauAnalysis-ShapeSmith.
source /cvmfs/sft.cern.ch/lcg/views/LCG_108/x86_64-el9-gcc15-opt/setup.sh # Python 3.12, uproot, pandas, mplhep, correctionlib, XRootD
python3 -m venv --system-site-packages ~/.venvs/shapesmith && source ~/.venvs/shapesmith/bin/activate
pip install -e ".[test]" # editable: changes in this checkout are live
pytest # tests on synthetic ntuples, no networkshapesmith validate -c run.yaml # build + validate the analysis, list processes and needed columns per channel
shapesmith skim -c run.yaml [--channels mt] [--samples TT,SingleMuon] [--force] [--workers 8] # ntuples (+ friends) -> Parquet
shapesmith hist -c run.yaml [--control] [--regions all] [--skip-systematics] [--processes ZTT,data] # histograms (ROOT + JSON index)
shapesmith estimate -c run.yaml [--control] # estimated processes and variations (Channel.estimators)
shapesmith measure -c run.yaml [--suggest-binning] [--merge] # Analysis.measurement -> <output_dir>/<measurement>/<era>/
shapesmith plot -c run.yaml [--control] [--region nominal] [--blind] [--log]
shapesmith publish -c run.yaml --to DIR --variant KEY [--label TEXT] [--description TEXT] [--channels et,mt] # plots + yields -> static web gallery DIR/index.html
shapesmith sync -c run.yaml # combine-style shape files per channel
shapesmith datacards -c run.yaml [--min-background 1.0] [--no-systematics]
shapesmith fit -c run.yaml [--final-states mt,all] [--skip-combine] # limits, significance, best fit
shapesmith ml-export -c run.yaml # Feather training folds
shapesmith inspect output/shapes.root [--unchanged] # what a histogram file contains; variations equal to their nominalskim reads the ntuples once into <skim_dir>/<channel>/<nick>/*.parquet plus a
manifest.json; everything else works on the skims. skim, hist, measure and ml-export
record what produced their outputs in <directory>/versions/<analysis name>.json (versions,
git hashes of core, analysis and sample database, the full run configuration and its hash).
publish copies the plots of <output_dir>/plots/ into a web gallery (DIR/index.html, a static
viewer reading manifest.json; also works from file://): one column per variant (--variant,
[a-z0-9_]+, e.g. one run configuration), with the data, prediction and stacked-group yields of
every plotted region and category (from the yield control variable where there is one, computed
as the stack plot does) and the provenance of the histograms. Publishing a variant again replaces
its published channels and keeps the other variants and channels; only changed plots are copied
and everything is world-readable. --title/--about set the gallery's title and the text of its
collapsed "About this gallery" section. A directory that is neither empty nor a gallery is refused.
web: {dir, variant, label, description, title, about} in the run YAML sets the defaults of the
options. For an analysis with a measurement, publish takes the measurement's plots from
<output_dir>/<measurement>/<era>/ instead (the fake factors: by process, split category and
quantity, with the p-value of each correction and the ttbar data/MC factors); a gallery holds
either control plots or one measurement.
Everything that differs by channel lives in its Channel: samples, cuts, processes, regions,
categories, control variables, variations and estimators.
Sample(nick, group, kind, xsec, nevents, generator_weight, cut): kinddata,mcorembedding;norm_weight = xsec / (nevents * generator_weight) * sign(genWeight)for MC. A sample's owncut(e.g. one generator-level part of an inclusive sample) is applied at skim time; its normalisation stays that of the whole sample.Process(name, group, role, plot_group, selection): roledata,signal,backgroundorauxiliary(booked in the nominal region only, for an estimator; never in datacards, plots or the ML export). A process owns its whole weight set.Region(name, replace_cuts, add_weights, replace_weights): cuts replaced by name, weights replaced where a process carries them, weights added to every process.WeightVariation(name, replace_weights, applies_to, regions)andColumnVariation(name, suffix | derived, applies_to, regions, groups): a column variation reads columncasc + suffixwhere that branch exists (a CROWN shift), or replaces it byderived[c], an expression of nominal columns;groupsrestricts it to the samples of some groups (a CROWN shift produced for some samples only). Names ending inUp/Downneed their partner and become datacard shapes; other names are templates.- Estimators, run in order by
estimate:DataMinus(output, region, subtract, scale, clip_negative),ABCD(output, b, c, d, subtract),TemplateShift(name, process, template, fraction), which also varies every template variationtof the process (as<name>Up@t), andVariationSum(name, parts):<name>Up/Down= nominal + the sum of (part - nominal) for every histogram with a part, the partpbeing the column variation<name>Up%p/<name>Down%p(histogram.part_of), e.g. one nuisance from CROWN shifts produced per decay mode (exact where no event is changed by two parts). Parts never become datacard shapes; place the sum after the estimators whose outputs should carry it.
The event rule (shapesmith.events, used by hist, the ML export and measurements): cuts are
the channel cuts with the region's replacements, the process cuts and the skim cuts; weights are
the process weights with the region's replacements, then the region's added weights; a weight
variation replaces weights in place (and does not apply to a process lacking one of them), a
column variation rewrites every expression. Event weight: (norm_weight * lumi) * product,
lumi for MC only.
Booking: the signal in the nominal region, other processes in the nominal and the estimator
regions (or --regions). A variation is filled where its applies_to and regions allow; a
column variation only where it changes an expression of the histogram (elsewhere it equals the
nominal). DataMinus builds its output for every column variation of data or a subtracted
process, taking the nominal of inputs that lack it; weight variations are not propagated, ABCD
is nominal only. --skip-systematics skips weight and column variations. hist always refills
the requested scopes and keeps the other histograms of its file.
Per sample, the skim keeps the columns of every expression its processes can use and the
shifted branches of its column variations. An event is kept if the skim cuts pass nominally or
under any column variation of its kind (hist re-applies the varied skim cuts). Where the channel
declares CROWN shifts for a sample (its kind and group), every file must carry a shifted branch for each of them,
and every shifted branch c__X of a needed column must be declared; otherwise the skim fails.
The manifest records the skim contract. A stored skim is reused when its skim cuts,
normalisation and sample cut are unchanged, its recorded column variations and friends contain
the current ones, and its Parquet schema holds the needed columns; so one skim made with all
variations and friends serves every subset. Incompatible skims are listed together, with the
skim --force --samples ... fix. A new friend version gets a new friend base (a friend rewritten
in place is not detected). The manifest is written before the first file of a sample and after
every finished file, so an interrupted skim resumes with only the unfinished files. Parquet
files and manifests are replaced atomically.
shapesmith.samples.read_sample_list reads a KingMaker sample list copied unchanged (one nick per
line); normalisation(database, nicks) looks every nick up in the KingMaker sample database
(datasets.json) and reports all missing nicks at once. The database renames nicks now and
then, so an older production needs the database checkout it was produced with.
Analysis.measurement is an object with a name and run(context); shapesmith measure calls
it with a MeasureContext (configuration, analysis, channels, output directory,
events(query), provenance(), and the flags suggest_binning and merge, the latter to combine
the results of earlier runs into one payload). Measurements:
shapesmith.measurements.tau_id_es: tau-ID scale factor and energy scale of embedded taus, the chain of smhtt_ultauID_SFs_dev. One run per working-point combination fills the mt channel and the mm control region, writes the shapes in the input format of MorphingTauID2017 and runs per category in CMSSW the datacards, T2W, a 2D likelihood scan, close-up 1D scans of both POIs and the MultiDimFit singles fit, the result (results.json, plots). As in the predecessor, the close-up scans and the singles fit start at the lowest grid point of the step before, within its 2 sigma intervals widened by a margin. The energy scale is a grid of template column variations in units of 0.1 % (grid.py);--mergewrites the correctionlib payload of all combinations. A 1 sigma interval at the fit range or a failed crossing is an error.shapesmith.measurements.fake_factors: the fake-factor measurement (the SM method of TauFakeFactors). An analysis setsAnalysis.measurement = FakeFactorMeasurement(legs), built from its region and process names and per-channel tables (Binned,Fit,Split,ProcessFF,DrSr,DataScale,Fractions,Leg).shapesmith measurewritesfake_factors_<ch>.json.gzandFF_corrections_<ch>.json.gz(the conventions of the CROWN fake-factor friend), the intermediate DR->SR payload,measurement_<ch>.json(every ratio, fit input and result) and plots;--suggest-binningprints equipopulated edges instead. The histograms (hist.py: centre of mass, MC-suppressed errors), fits (fit.py: bands, SystMCShift, the compatibility-with-1 reset, sparsify) and the kernel are checked against ROOT and TauFakeFactors (tests/reference/make_fake_factor_reference.py); the normative algorithm is the "Algorithm reference" of the design spec.
Building blocks:
shapesmith.payloads: correctionlib schema-v2 builders; the provenance is JSON inCorrectionSet.description; payloads are checked (schema, evaluator, no$schemakey, so correctionlib 2.6 reads them) and written gzip compressed.shapesmith.measurements.smoothing: TauFakeFactors' kernel smoothing without ROOT, a port of ROOT'sTGraphSmooth::SmoothKernchecked against ROOT and FF_Updated (tests/reference/make_smoothing_reference.py).shapesmith.cmssw: run commands in a CMSSW environment (combine.cmssw_dir), started from a clean environment so that the Python environment of ShapeSmith does not shadow CMSSW's.
analysis: my_analysis.analysis:build # module:function returning an Analysis
era: "2018"
channels: [et, mt, tt]
switches: {jet_fakes: mc, embedding: false} # free-form, interpreted by the analysis
sample_database: ../../KingMaker_sample_database/nanoAOD_v15/datasets.json # relative paths: relative to this file
ntuples:
server: root://cmsdcache-kit-disk.gridka.de
base: /store/user/USER/CROWN/ntuples/TAG/CROWNRun
friends:
- {base: /store/user/USER/CROWN/ntuples/TAG/CROWNFriends/nn_v1, applies_to: [mc, data, embedding]}
skim_dir: /ceph/USER/shapesmith/TAG
output_dir: output/TAG
ml_dir: /ceph/USER/shapesmith_ml/TAG
workers: 16
log_level: INFO # DEBUG | INFO | WARNING | ERROR
combine: {cmssw_dir: /work/USER/CMSSW_14_1_9, scram_arch: el9_amd64_gcc12}
web: {dir: /web/USER/public_html/TAG, variant: emb_ff, label: "embedding + FF"} # shapesmith publishlog_level sets the level of ShapeSmith and of the analysis package; one run changes it with
-s log_level=DEBUG. Other libraries (matplotlib, numexpr, uproot, ...) log from WARNING on.
- Console: the messages, on stderr (stdout keeps the reports of
validateandinspect). - Log file: every command but
validateandinspectalso writes<output_dir>/logs/<command>_<YYYYmmdd-HHMMSS>.logwith time, level, process and logger, and the traceback of a failing command. - INFO: the run (versions, configuration, channels, workers), the analysis per channel, what every stage does per channel, the progress of the parallel jobs (every tenth), the files written and the duration. DEBUG adds every ntuple file, booking, estimate, datacard bin, fit and plot, and the output of the CMSSW commands (at INFO it is shown only when they fail).
- The worker processes log through the main process, so their messages reach the console and the log file as well.
- Expressions (cuts, weights, variables) are pandas
evalsyntax on ntuple columns, evaluated with numexpr:(pt_1 > 25) & (abs(eta_1) < 2.1),id_wgt_tau_1 * trg_wgt. - The edges of a variable are numbers or
EqualData(n_bins, low, high):n_binsbins with equal data counts in (low, high), as smhtt_ulgof/build_binning.py(binning.py). Every fill (hist, a measurement) computes them from the data of the category in the nominal region, uses them for every region and variation, and records them inbinning.jsonnext to the histogram file. - Histograms are the small numpy
shapesmith.histogram.Histogram; ROOT files are written and read through uproot (boost-histogram/histare broken in LCG_108, their axis edges come out constant). workers > 1uses a spawn-based process pool (fork deadlocks with uproot/pyarrow threads). Scripts that call ShapeSmith functions with more than one worker therefore need the usualif __name__ == "__main__":guard;--workers 1runs everything inline for debugging.tests/golden/holds the mini-dataset outputs of coreaa5daa5; the rebuilt core reproduces them bitwise (tests/test_golden.py).