RiT: Vanilla Diffusion Transformers Suffice in Representation Space
TL;DR AI
2 min readKey summary
RiT-XL checkpoint, weights, and evaluation code have been released for ImageNet 256×256 generation in representation space.
The model uses frozen DINOv2-Small features and standard sampling, without extra distillation.
It reports strong FID results, including 1.45 at CFG=1 and 1.14 at CFG≈3.7 with 25 Heun steps.
These numbers beat several prior diffusion transformer variants, showing a vanilla diffusion transformer can perform very well with a frozen encoder.
