One cutlist. Everything derived from it. Every sync claim measured.
Cutting a documentary means reconciling sources that don't agree: a camera and a field recorder that were never locked together, a transcript whose word timings are slightly wrong, a paper edit that drifts from the timeline the moment someone trims a beat.
cutlist holds those together with two rules.
One cutlist, everything derived from it. The paper edit, the draft render, the NLE project, the subtitles and the verification all read the same beats. They cannot drift apart, because there is only one set of timecodes and nothing downstream is allowed to invent one.
Measure, don't assume — and run the control first. Every sync claim is measured against source. When a measurement can't be made confidently, it says so instead of returning a number.
pip install cutlist # ffmpeg must be on PATHgit clone https://github.com/tekgrunt/cutlist && cd cutlist
python examples/synthetic/make_fixture.py # a fake shoot with known answers
cd examples/synthetic
cutlist ingest && cutlist sync && cutlist refine
cutlist export draft && cutlist verifyThe fixture's field recorder is deliberately 12.5 s early and 40 ppm fast,
and its transcript ends every word 180 ms early. cutlist sync recovers the
offset to under a millisecond and the clock ratio to about 1 ppm. cutlist refine
finds every clipped word ending.
Those aren't smoke tests — the generator knows the right answers, so the suite asserts against truth rather than against it ran.
flowchart LR
A["project.yaml<br/>the shoot"] --> C
B["edit.yaml<br/>the paper edit"] --> C
C["cutlist.json<br/>ONE SOURCE OF TRUTH"]
C --> D["draft mp4"]
C --> E["Kdenlive / FCPXML"]
C --> F["SRT + VTT"]
C --> G["script page"]
D & E & F & G --> H["verify<br/>proves they all agree"]
| command | what it does |
|---|---|
cutlist init |
write a starting project.yaml, scanning a media directory |
cutlist ingest |
probe sources, extract 16 kHz working audio |
cutlist sync |
align recorders and second angles; measure drift |
cutlist transcribe |
word-level transcripts (faster-whisper) |
cutlist diarize |
label segments by voice fingerprint or mouth motion |
cutlist find |
frame-accurate in/out for a spoken phrase |
cutlist refine |
snap every cut to silence so no word is clipped |
cutlist tighten |
cut to a target runtime, using the notes as instructions |
cutlist audio |
measure what separates sources; derive the EQ from them |
cutlist export |
draft · kdenlive · fcpxml · subs · paperedit |
cutlist verify |
prove the edit is what the cutlist says |
cutlist status |
what exists, what is stale, what is next |
The camera runs at one sample rate, the recorders at another, and nothing is timecode-locked. Two separate things can be wrong:
- offset — the recorder was rolling before or after the camera. One number.
- drift — the clocks don't tick at the same rate, so the correct offset at minute 40 isn't the offset at minute 0.
So cutlist sync probes at several points across each recording and fits a line.
The slope is the clock ratio; the intercept is the offset. Sync on a single
point and the far end of a long take slides out of lip sync.
Which recording belongs to which camera is decided the same way — by whichever correlates. Filenames lie.
$ cutlist sync
===== RECORDER (recorder) =====
probe @ 69.3s -> CAM_A 81.758s z= 16.3
probe @ 128.2s -> CAM_A 140.756s z= 16.2
probe @ 187.2s -> CAM_A 199.753s z= 17.4
match confidence: CAM_A z=17, CAM_B z=4
offset at start : +12.501 s
clock ratio : 1.00003962 (+40 ppm)
drift over take : +10 ms across 4.3 min
-> a single offset is enough, no resampling neededWhisper's word-level end times land early — they mark where the model decided the word was recognisable, not where the sound stops. Cut on them and you chop the final consonant. Measured on real material: whisper's OUT sat at −16.3 dB, mid-word, and the word didn't decay for another 230 ms.
cutlist refine keeps the transcript's timings as the intent, then moves each
boundary to the nearest real silence, bounded so a neighbouring word is never
swallowed. Then it snaps both onto the source frame grid — with -ss, ffmpeg
starts audio at the exact sample but video at the first frame at or after it, so
a cut landing mid-frame leaves audio up to a frame ahead of picture.
Because two earlier measurements on the project this came from were confidently wrong, and shipped.
An earlier lip-sync check compared the rendered soundtrack against the camera at a timeline position computed by adding up beat durations. Every encoded segment rounds up to a whole frame, so the position model drifted a few ms per segment and reported ~180 ms of "sync error" by the end of the film — all of it the test's own arithmetic.
So cutlist verify assumes nothing about position. For a moment in the render it
establishes two things independently:
- where the audio came from — correlate the rendered sound against source
- where the picture came from — match the rendered frame against source frames
If those disagree, the difference is real. And the control runs first, so figures are read against what no error looks like for that measurement, not against zero.
$ cutlist verify
[PASS] runtime: model 26.326s vs render 26.352s (diff 26 ms, tolerance 133 ms)
[PASS] filters: render chain 'volume=0dB' shifts audio by +0.0 ms
[PASS] sync @5s: audio 5.635s vs picture 5.635s -> +0 ms (z=9, margin 0.07)
[SKIP] sync @10s: could not localise the audio (z=6, 3.7s window) — beat too
short or too quiet to measure, not an error
3 passed, 0 failed, 3 not measurableChecks have three outcomes, not two. SKIP means this wasn't measurable here
— a beat too short to correlate, a frame too static to localise. Only FAIL means
measured, and wrong. Conflating them is how a verification layer starts crying
wolf and gets ignored.
That filter check isn't hypothetical: ffmpeg's alimiter delays audio by its
attack window, about 5 ms, which reads as lip-sync error on every cut that uses
it. The check caught it in this project's own renderer.
acts:
- title: Cold open
blurb: Open on doubt, resolve it in ten seconds.
beats:
- clip: C0091
in: '00:28:40.210'
out: '00:28:51.530'
speaker: Max
quote: >-
There's another famous saying, which is: being early is the same as
being wrong.
note: The hook. State the doubt before anyone else can.
trim: 28A note isn't decoration. cutlist tighten reads quoted phrases inside a note
as anchors and won't trim them away:
Trim to the waterfall/DevOps contrast. The
'everybody stay late'detail is the keeper.
It enumerates the windows that begin and end on a sentence boundary and keeps the
one nearest the target that still contains the anchors. Cutting mid-sentence is
never a candidate. If the film is still long, whole beats come out rather than
every beat being shaved — a shorter film of complete thoughts beats a film of
clipped ones — in the order the notes imply: optional, then cut_first, never
protect.
A locked-off two-shot on one mono mic has nothing to separate voices with. On the source project the L/R channels correlated at 0.9998 — one mic, duplicated.
Two ways out, and cutlist diarize picks per source:
- voice — build a fingerprint per speaker from the source where they're the dominant voice, then label everything by nearest match. In a solo interview whoever talks more is the subject, because the interviewer only asks questions — so the reference needs no hand-labelling.
- motion — whoever's mouth is moving is talking. One sequential pass over the video, recording motion energy per speaker's mouth region.
The core needs only numpy, PyYAML and ffmpeg. The expensive dependencies are each confined to one stage, so you install them only if you use that stage:
pip install cutlist # core
pip install 'cutlist[transcribe]' # + faster-whisper
pip install 'cutlist[diarize]' # + speechbrain, torch, opencvYou can skip both. Write transcripts/<id>.json yourself with
segments[].words[] carrying start / end / word, and set
segments[].speaker.
pip install -e '.[dev]'
python examples/synthetic/make_fixture.py
pytest -qThe fixture media and its edit.yaml are generated rather than committed.
That's deliberate: edit.yaml was committed once, went stale when the generator
changed, and its beats silently pointed at audio that no longer existed in that
form. Deriving it is the same rule the tool runs on.
If you're changing anything in sync, refine or timecode, the tests in
tests/test_pipeline.py assert against the generator's known values — keep them
that way.
Extracted from the tooling built to cut a founder documentary: two subjects, four camera files, three field recorders, nothing timecode-locked. The comments in this code that explain why something is done a particular way are mostly scar tissue from that shoot, and they were kept deliberately.
MIT — see LICENSE.