Skip to content

Latest commit

 

History

3 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

cutlist

Build a documentary from footage whose clocks disagree.

One cutlist. Everything derived from it. Every sync claim measured.

tests PyPI Python License


Cutting a documentary means reconciling sources that don't agree: a camera and a field recorder that were never locked together, a transcript whose word timings are slightly wrong, a paper edit that drifts from the timeline the moment someone trims a beat.

cutlist holds those together with two rules.

One cutlist, everything derived from it. The paper edit, the draft render, the NLE project, the subtitles and the verification all read the same beats. They cannot drift apart, because there is only one set of timecodes and nothing downstream is allowed to invent one.

Measure, don't assume — and run the control first. Every sync claim is measured against source. When a measurement can't be made confidently, it says so instead of returning a number.

pip install cutlist        # ffmpeg must be on PATH

Try it in 60 seconds, without footage

git clone https://github.com/tekgrunt/cutlist && cd cutlist
python examples/synthetic/make_fixture.py     # a fake shoot with known answers

cd examples/synthetic
cutlist ingest && cutlist sync && cutlist refine
cutlist export draft && cutlist verify

The fixture's field recorder is deliberately 12.5 s early and 40 ppm fast, and its transcript ends every word 180 ms early. cutlist sync recovers the offset to under a millisecond and the clock ratio to about 1 ppm. cutlist refine finds every clipped word ending.

Those aren't smoke tests — the generator knows the right answers, so the suite asserts against truth rather than against it ran.

The pipeline

flowchart LR
  A["project.yaml<br/>the shoot"] --> C
  B["edit.yaml<br/>the paper edit"] --> C
  C["cutlist.json<br/>ONE SOURCE OF TRUTH"]
  C --> D["draft mp4"]
  C --> E["Kdenlive / FCPXML"]
  C --> F["SRT + VTT"]
  C --> G["script page"]
  D & E & F & G --> H["verify<br/>proves they all agree"]
Loading
command what it does
cutlist init write a starting project.yaml, scanning a media directory
cutlist ingest probe sources, extract 16 kHz working audio
cutlist sync align recorders and second angles; measure drift
cutlist transcribe word-level transcripts (faster-whisper)
cutlist diarize label segments by voice fingerprint or mouth motion
cutlist find frame-accurate in/out for a spoken phrase
cutlist refine snap every cut to silence so no word is clipped
cutlist tighten cut to a target runtime, using the notes as instructions
cutlist audio measure what separates sources; derive the EQ from them
cutlist export draft · kdenlive · fcpxml · subs · paperedit
cutlist verify prove the edit is what the cutlist says
cutlist status what exists, what is stale, what is next

Why the sync works this way

The camera runs at one sample rate, the recorders at another, and nothing is timecode-locked. Two separate things can be wrong:

  • offset — the recorder was rolling before or after the camera. One number.
  • drift — the clocks don't tick at the same rate, so the correct offset at minute 40 isn't the offset at minute 0.

So cutlist sync probes at several points across each recording and fits a line. The slope is the clock ratio; the intercept is the offset. Sync on a single point and the far end of a long take slides out of lip sync.

Which recording belongs to which camera is decided the same way — by whichever correlates. Filenames lie.

$ cutlist sync
===== RECORDER (recorder) =====
    probe @    69.3s -> CAM_A    81.758s  z= 16.3
    probe @   128.2s -> CAM_A   140.756s  z= 16.2
    probe @   187.2s -> CAM_A   199.753s  z= 17.4
  match confidence: CAM_A z=17, CAM_B z=4
  offset at start : +12.501 s
  clock ratio     : 1.00003962  (+40 ppm)
  drift over take : +10 ms across 4.3 min
  -> a single offset is enough, no resampling needed

Why cuts get snapped to silence

Whisper's word-level end times land early — they mark where the model decided the word was recognisable, not where the sound stops. Cut on them and you chop the final consonant. Measured on real material: whisper's OUT sat at −16.3 dB, mid-word, and the word didn't decay for another 230 ms.

cutlist refine keeps the transcript's timings as the intent, then moves each boundary to the nearest real silence, bounded so a neighbouring word is never swallowed. Then it snaps both onto the source frame grid — with -ss, ffmpeg starts audio at the exact sample but video at the first frame at or after it, so a cut landing mid-frame leaves audio up to a frame ahead of picture.

Why there is a verify step

Because two earlier measurements on the project this came from were confidently wrong, and shipped.

An earlier lip-sync check compared the rendered soundtrack against the camera at a timeline position computed by adding up beat durations. Every encoded segment rounds up to a whole frame, so the position model drifted a few ms per segment and reported ~180 ms of "sync error" by the end of the film — all of it the test's own arithmetic.

So cutlist verify assumes nothing about position. For a moment in the render it establishes two things independently:

  • where the audio came from — correlate the rendered sound against source
  • where the picture came from — match the rendered frame against source frames

If those disagree, the difference is real. And the control runs first, so figures are read against what no error looks like for that measurement, not against zero.

$ cutlist verify
  [PASS] runtime: model 26.326s vs render 26.352s (diff 26 ms, tolerance 133 ms)
  [PASS] filters: render chain 'volume=0dB' shifts audio by +0.0 ms
  [PASS] sync @5s: audio 5.635s vs picture 5.635s -> +0 ms (z=9, margin 0.07)
  [SKIP] sync @10s: could not localise the audio (z=6, 3.7s window) — beat too
         short or too quiet to measure, not an error

3 passed, 0 failed, 3 not measurable

Checks have three outcomes, not two. SKIP means this wasn't measurable here — a beat too short to correlate, a frame too static to localise. Only FAIL means measured, and wrong. Conflating them is how a verification layer starts crying wolf and gets ignored.

That filter check isn't hypothetical: ffmpeg's alimiter delays audio by its attack window, about 5 ms, which reads as lip-sync error on every cut that uses it. The check caught it in this project's own renderer.

The paper edit is data

acts:
  - title: Cold open
    blurb: Open on doubt, resolve it in ten seconds.
    beats:
      - clip: C0091
        in:  '00:28:40.210'
        out: '00:28:51.530'
        speaker: Max
        quote: >-
          There's another famous saying, which is: being early is the same as
          being wrong.
        note: The hook. State the doubt before anyone else can.
        trim: 28

A note isn't decoration. cutlist tighten reads quoted phrases inside a note as anchors and won't trim them away:

Trim to the waterfall/DevOps contrast. The 'everybody stay late' detail is the keeper.

It enumerates the windows that begin and end on a sentence boundary and keeps the one nearest the target that still contains the anchors. Cutting mid-sentence is never a candidate. If the film is still long, whole beats come out rather than every beat being shaved — a shorter film of complete thoughts beats a film of clipped ones — in the order the notes imply: optional, then cut_first, never protect.

Diarization with nothing spatial to work with

A locked-off two-shot on one mono mic has nothing to separate voices with. On the source project the L/R channels correlated at 0.9998 — one mic, duplicated.

Two ways out, and cutlist diarize picks per source:

  • voice — build a fingerprint per speaker from the source where they're the dominant voice, then label everything by nearest match. In a solo interview whoever talks more is the subject, because the interviewer only asks questions — so the reference needs no hand-labelling.
  • motion — whoever's mouth is moving is talking. One sequential pass over the video, recording motion energy per speaker's mouth region.

Installing

The core needs only numpy, PyYAML and ffmpeg. The expensive dependencies are each confined to one stage, so you install them only if you use that stage:

pip install cutlist                 # core
pip install 'cutlist[transcribe]'   # + faster-whisper
pip install 'cutlist[diarize]'      # + speechbrain, torch, opencv

You can skip both. Write transcripts/<id>.json yourself with segments[].words[] carrying start / end / word, and set segments[].speaker.

Contributing

pip install -e '.[dev]'
python examples/synthetic/make_fixture.py
pytest -q

The fixture media and its edit.yaml are generated rather than committed. That's deliberate: edit.yaml was committed once, went stale when the generator changed, and its beats silently pointed at audio that no longer existed in that form. Deriving it is the same rule the tool runs on.

If you're changing anything in sync, refine or timecode, the tests in tests/test_pipeline.py assert against the generator's known values — keep them that way.

Credits

Extracted from the tooling built to cut a founder documentary: two subjects, four camera files, three field recorders, nothing timecode-locked. The comments in this code that explain why something is done a particular way are mostly scar tissue from that shoot, and they were kept deliberately.

MIT — see LICENSE.

About

Build a documentary from footage whose clocks disagree. One cutlist, everything derived from it; every sync claim measured.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages