DistillAlign: Coordinating Mode Covering and Mode Seeking in Autoregressive Video Distillation
TL;DR AI
2 min readKey summary
Researchers reexamined autoregressive video distillation through a distributional lens and found that standard DMD training can produce students with high precision but poor coverage.
They also showed that late-stage training can further collapse diversity, making students more mode-seeking and less representative of the teacher distribution.
To address this, they introduced DistillAlign, a shared-latent evaluation protocol and a joint distillation objective that combines DMD with Consistency Distillation.
The method improved video generation quality, coverage, diversity, and refinement stability on benchmarks such as VBench and models including Wan-1.3B and Wan-14B.
