Switch language한국어
Back to the list

DistillAlign: Coordinating Mode Covering and Mode Seeking in Autoregressive Video Distillation

TL;DR AI

Key summary

2 min read
  1. Researchers reexamined autoregressive video distillation through a distributional lens and found that standard DMD training can produce students with high precision but poor coverage.

  2. They also showed that late-stage training can further collapse diversity, making students more mode-seeking and less representative of the teacher distribution.

  3. To address this, they introduced DistillAlign, a shared-latent evaluation protocol and a joint distillation objective that combines DMD with Consistency Distillation.

  4. The method improved video generation quality, coverage, diversity, and refinement stability on benchmarks such as VBench and models including Wan-1.3B and Wan-14B.

Read the original