Switch language한국어
Back to the list

Accelerating Text-to-Video Generation with Calibrated Sparse Attention

TL;DR AI

Key summary

2 min read
  1. Researchers introduced CalibAtt, a training-free method for accelerating text-to-video diffusion models.

  2. It uses an offline calibration step to identify stable sparse and repeated attention patterns, then skips low-value connections during inference.

  3. On models like Wan 2.1 14B and Mochi 1, it speeds up generation by up to 1.58× while preserving quality.

  4. The method makes high-resolution and few-step video generation more practical without retraining.

Read the original