Accelerating Text-to-Video Generation with Calibrated Sparse Attention

TL;DR AI
2 min readKey summary
Researchers introduced CalibAtt, a training-free method for accelerating text-to-video diffusion models.
It uses an offline calibration step to identify stable sparse and repeated attention patterns, then skips low-value connections during inference.
On models like Wan 2.1 14B and Mochi 1, it speeds up generation by up to 1.58× while preserving quality.
The method makes high-resolution and few-step video generation more practical without retraining.
