Skip to content

Repository files navigation

Coevolving Representations in
Joint Image-Feature Diffusion

Theodoros Kouzelis1,3   ·   Spyros Gidaris2   ·   Nikos Komodakis1,4,5
1 Archimedes/Athena RC   2 valeo.ai   3 National Technical University of Athens  
4 University of Crete   5 IACM-Forth  

📃 Paper   🔗 ReDi (our previous work)  

teaser.png

CoReDi builds on our previous work, ReDi.

Pretrained models will be released soon. Stay tuned!

Setup

The environment is managed with uv (Python 3.11, PyTorch 2.8 + CUDA 12.8):

uv sync --extra eval      # --extra eval installs TensorFlow, needed only for FID evaluation
source .venv/bin/activate

Training also downloads stabilityai/sd-vae-ft-ema (Hugging Face) and DINOv2 (torch.hub) on first use.

Data preparation

We use the same ImageNet preprocessing as REPA: images are center-cropped to 256x256 and encoded with the SD-VAE.

Training

Edit the paths in train.sh and run bash train.sh, or launch it directly:

torchrun --nnodes 1 --nproc_per_node 8 train_coredi.py \
   --model "SiT-B/2" \
   --feature-path [TARGET_PATH] \
   --exp-name "SiT-B-2-coredi" \
   --results-dir results \
   --global-batch-size 256 \
   --sg-dino-loss True \
   --projection-bn 1d \
   --dino-proj-lr 1e-4 \
   --vic-loss batch \
   --vic-lambda-var-features 1 \
   --vic-lambda-cov 0 \
   --random-init \
   --log-pca-iter 5000 \
   --max-train-steps 200000 \
   --no-wandb

Checkpoints are saved to results/<exp-name>/checkpoints/ every --ckpt-every steps (default 50k). Resume with --ckpt results/<exp-name>/checkpoints/<step>.pt. Drop --no-wandb to log to Weights & Biases.

Sampling and evaluation

Download the ADM reference batch for ImageNet 256x256 (VIRTUAL_imagenet256_labeled.npz) into the repo root. Then edit the paths in sample.sh and run bash sample.sh, or:

torchrun --nnodes 1 --nproc_per_node 8 sample_ddp.py SDE \
    --model "SiT-B/2" \
    --pca-rank 8 \
    --ckpt results/SiT-B-2-coredi/checkpoints/0200000.pt \
    --cfg-scale 1.0 \
    --num-fid-samples 50000 \
    --per-proc-batch-size 128 \
    --num-sampling-steps 250 \
    --ref-batch VIRTUAL_imagenet256_labeled.npz \
    --sample-dir results/SiT-B-2-coredi/

This generates 50k images, packs them into an .npz and computes FID, sFID, IS, precision and recall. Metrics are written to results/SiT-B-2-coredi/fids/. Use ODE instead of SDE for the deterministic sampler. sample.sh also sets LD_LIBRARY_PATH so the TensorFlow evaluator can use the GPU.

Citation

If you found CoReDi useful in your research, please consider starring ⭐ us on GitHub and citing 📚 us in your research!

@article{kouzelis_coredi2026,
  title={Coevolving Representations in Joint Image-Feature Diffusion},
  author={Kouzelis, Theodoros and Gidaris, Spyros and Komodakis, Nikos},
  journal={arXiv preprint arXiv:2604.17492},
  year={2026}
}

About

[ECCV'26] Coevolving Representations in Joint Image-Feature Diffusion

Resources

Stars

15 stars

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages