Switch language한국어
Back to the list

SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation

TL;DR AI

Key summary

2 min read
  1. Researchers unveiled SANA-Video 2.0, a hybrid video diffusion transformer for efficient video generation.

  2. The 5B and 14B models combine linear and softmax attention, plus attention residuals, to improve token interaction.

  3. Trained from scratch, the system can generate up to 720p video efficiently on a single GPU.

  4. The approach aims to keep much of softmax attention’s quality while preserving the scalability of linear attention.

Read the original