Turn a long YouTube video (or local file) into short, captioned, vertical clips — fully local, no paid APIs required. Good fit for faceless-video channels: no face tracking, just a blurred-background 9:16 reframe + clean word-by-word captions.
- Download the source video with
yt-dlp(or point it at a local file). - Transcribe locally with
faster-whisper, getting word-level timestamps. - Score the transcript with text heuristics (hooks, questions, numbers, superlatives, filler-word penalty) to find the most clip-worthy moments — no LLM API needed.
- Cut & caption:
ffmpegtrims each clip, reframes it to a blurred- background 9:16 vertical layout, and burns in animated word-by-word captions.
# 1. System dependency: ffmpeg (not installable via pip)
# macOS: brew install ffmpeg
# Ubuntu: sudo apt install ffmpeg
# Windows: choco install ffmpeg (or download from ffmpeg.org)
# 2. Python dependencies
pip install -r requirements.txtFirst run will download the Whisper model (a few hundred MB, one-time).
# From a YouTube URL
python main.py --url "https://www.youtube.com/watch?v=XXXXXXXX" --num-clips 5
# From a file you already have
python main.py --file /path/to/podcast.mp4 --num-clips 8
# Keep the original 16:9 aspect ratio instead of vertical reframe
python main.py --url "..." --horizontalOutput clips land in ./output/clips/clip_01.mp4, clip_02.mp4, etc.
Everything lives in config.py:
WHISPER_MODEL— bigger = more accurate, slower (base→small.en→medium.en)MIN_CLIP_DURATION/MAX_CLIP_DURATION— target clip lengthCAPTION_MAX_WORDS_ON_SCREEN— how many words shown per caption chunk- Caption colors/fonts
The heuristic scorer in highlights.py works with zero setup, but it's not
as smart as an LLM. If you install Ollama and pull a
model (ollama pull llama3.1), you can swap in rank_with_ollama() from
highlights.py to have a local LLM score each candidate clip instead —
still completely free, just needs a decent GPU/CPU to run at reasonable speed.
- Whisper runs on CPU by default (
compute_type="int8") so it works everywhere, but it's slow on long videos. If you have an NVIDIA GPU with CUDA, edittranscribe.pyto usedevice="cuda", compute_type="float16"for a large speedup. - This tool doesn't do face tracking/cropping — it's built for faceless content (podcasts, voiceovers, screen recordings) where a centered, blurred-background vertical layout looks clean without needing to track a speaker's face.