Chimera: Designing and Chinchilla-Scaling Hybrid Visual Diffusion Transformers
TL;DR AI
2 min readKey summary
Researchers introduced Chimera, a hybrid diffusion backbone for text, image, and video tokens.
It combines KDA, MLA, short convolutions, and sparse MoE layers to improve efficiency for long-context generation.
They also propose HeteroP, a scaling recipe that tunes width and depth, and train an 11B model with 2B activated parameters.
Experiments show better compute efficiency than a full-attention baseline and strong zero-shot video length extrapolation.
