DistillAlign: Coordinating Mode Covering and Mode Seeking in Autoregressive Video Distillation
Anonymous release for ICLR 2027 submission
DistillAlign aligns and balances the mode-covering and mode-seeking objectives of the multi-stage video distillation pipeline — using only a 1.3B DMD teacher, it already surpasses baselines refined with a 14B DMD teacher.
Checkpoints
All released generators are Wan2.1-1.3B students; the size in each name refers to the teacher used during training.
| Model | Checkpoint | Description |
|---|---|---|
| Initializer (1.3B teacher) | distillalign_init_1p3b_teacher.pt | Pre-DMD initializer, trained toward a Wan2.1-T2V-1.3B teacher |
| Initializer (14B teacher) | distillalign_init_14b_teacher.pt | Pre-DMD initializer, trained toward a Wan2.1-T2V-14B teacher |
| Distilled (1.3B teacher) | distillalign_distill_1p3b_teacher.pt | Final joint-distilled generator, Wan2.1-T2V-1.3B DMD teacher |
| Distilled (14B teacher) | distillalign_distill_14b_teacher.pt | Final joint-distilled generator, Wan2.1-T2V-14B DMD teacher |
Each checkpoint stores a plain {"generator": state_dict} and loads directly
with the inference and evaluation code in the supplementary material.
Teacher Reference Caches
Ready-made teacher reference features for the teacher-normalized
distribution evaluation. Pass the .npz file to --teacher-features and
only the student side needs to be generated.
| Cache | Download | Description |
|---|---|---|
| Wan2.1-1.3B teacher reference features | wan2.1_t2v_1.3b_reference_vjepa2.npz | 256 x 2560 V-JEPA2 features, ready for --teacher-features |
| Wan2.1-14B teacher reference features | wan2.1_t2v_14b_reference_vjepa2.npz | 256 x 2560 V-JEPA2 features, ready for --teacher-features |
See teacher_caches/README.md for the exact sampling and extraction protocol.