This repository is a compact code package for the DyStream assessment work. It is not a full copy of the original DyStream repository. It contains the files that were added or modified for:
- Two-sample overfit training with
train.py. - Official-checkpoint, audio-overfit, and text-overfit inference.
- Text-condition injection where
captionis used as the prompt during training and inference.
The full file-by-file explanation is in:
docs/overfit_text_code_map_20260618.md
Important files:
train.py
main.py
video_to_latent.py
datasets/single_dyadic_prev_audio.py
model/motion_generation/text_conditioned_audio2face.py
configs/motion_gen/overfit_2samples_selected.yaml
configs/motion_gen/overfit_2samples_selected_text.yaml
scripts/extract_video_motion_latents.py
scripts/prepare_overfit_metadata.py
scripts/prepare_selected_overfit_samples.sh
scripts/run_selected_overfit.sh
scripts/run_official_overfit_infer.sh
scripts/run_finetuned_infer.sh
scripts/run_text_finetuned_infer.sh
data_json/overfit_items_selected.json
data_json/overfit_train_selected.json
data_json/overfit_test_selected.json
data_overfit_selected/sample_001/
data_overfit_selected/sample_002/
Large runtime assets are intentionally not included:
checkpoints/
tools/
generated videos
raw source videos
The original DyStream repository, official checkpoint, and wrapping encoder/decoder assets are still required to run this code.
The overfit experiment uses two manually selected talking-head clips. The selected train metadata is expanded into many sliding windows, instead of using only two metadata rows. This is important because DyStream trains on a fixed-length window; if the train JSON only has two rows, training only sees the beginning of each video.
The text-condition extension adds a text-conditioned model variant. Captions are read from metadata, passed through the dataset and training/inference code, encoded with a text encoder, and injected into the generation hidden states.
The two selected overfit samples are checked in under:
data_overfit_selected/sample_001/
data_overfit_selected/sample_002/
Each sample includes:
audio.wav
gt.mp4
motion.npz
preview.jpg
ref.png
ref_resize.png
Run inside a prepared DyStream checkout:
cd /ICML/ZJU/DyStream-mainPrepare selected metadata and motion latents:
TRAIN_STRIDE_FRAMES=1 DEVICE=cuda BATCH_SIZE=64 bash scripts/prepare_selected_overfit_samples.shIf you already have a cropped gt.mp4 and only need to rebuild the DyStream motion latent file, use either entry point below. They are equivalent:
python video_to_latent.py \
--video data_overfit_selected/sample_001/gt.mp4 \
--output data_overfit_selected/sample_001/motion.npz \
--device cuda \
--batch-size 64
python scripts/extract_video_motion_latents.py \
--video data_overfit_selected/sample_001/gt.mp4 \
--output data_overfit_selected/sample_001/motion.npz \
--device cuda \
--batch-size 64This requires the original DyStream tools/visualization_0416 assets and their wrapping/motion encoder checkpoints to be present in the checkout.
Pure audio overfit:
GPU_ID=0 \
TRAIN_BS=64 \
VAL_BS=2 \
MAX_STEPS=8000 \
CONFIG=configs/motion_gen/overfit_2samples_selected.yaml \
EXP_NAME=overfit_2samples_selected_stride1_bs64 \
bash scripts/run_selected_overfit.shText-conditioned overfit:
GPU_ID=0 \
TRAIN_BS=64 \
VAL_BS=2 \
MAX_STEPS=8000 \
CONFIG=configs/motion_gen/overfit_2samples_selected_text.yaml \
EXP_NAME=overfit_2samples_selected_text_stride1_bs64 \
bash scripts/run_selected_overfit.shOfficial checkpoint inference:
GPU_ID=0 \
CONFIG=configs/motion_gen/overfit_2samples_selected.yaml \
EXP_NAME=official_selected_infer_package \
DENOISING_STEPS=10 \
bash scripts/run_official_overfit_infer.shAudio-overfit inference:
GPU_ID=0 \
CONFIG=configs/motion_gen/overfit_2samples_selected.yaml \
CKPT=checkpoints/package_audio_step1500.ckpt \
EXP_NAME=audio_selected_step1500_infer_package \
DENOISING_STEPS=10 \
bash scripts/run_finetuned_infer.shText-overfit inference:
GPU_ID=0 \
CONFIG=configs/motion_gen/overfit_2samples_selected_text.yaml \
CKPT=checkpoints/package_text_step1500.ckpt \
EXP_NAME=text_selected_step1500_infer_package \
DENOISING_STEPS=10 \
bash scripts/run_text_finetuned_infer.shThe local generated comparison package used this structure:
deliverables/selected_comparison_step1500/
01_gt/
02_official_infer/
03_audio_overfit_infer/
04_text_overfit_infer/
Each folder contains two MP4 files, one for each selected sample.
resume_mode=weights_onlyloads official model weights without restoring the original trainer step or optimizer state.- The text-conditioned implementation satisfies the requirement that
captionis used as prompt during training. It is a first-pass conditioning implementation, not a mature strong-control model. - If a checkpoint filename contains
=, create a clean symlink before passing it through config overrides, for examplecheckpoints/package_audio_step1500.ckpt.
This repository also includes the non-weight files for the lagged realtime microphone-to-MP4 command collected from /root/autodl-tmp/DyStream_cudagraph_streamtest.
Run it with:
bash scripts/run_realtime_mic_to_mp4.shThe helper defaults to the requested command:
OMP_NUM_THREADS=1 TRANSFORMERS_OFFLINE=1 HF_HUB_OFFLINE=1 \
CUDA_VISIBLE_DEVICES=0,1 python -u realtime_mic_to_mp4.py \
--feature_lag_frames 3 \
--hop_ms 200 \
--denoising_steps 1 \
--motion_gpu 0 \
--render_gpu 1 \
--port 6008Model weights are intentionally excluded from git. See docs/realtime_mic_to_mp4.md for the required local asset paths.