RayDer: Scalable Self-Supervised Novel View Synthesis from Real-World Video

TL;DR AI
2 min readKey summary
Researchers introduced RayDer, a unified feed-forward transformer for self-supervised novel view synthesis from unconstrained video.
It combines camera estimation, reconstruction, and rendering in one model, using a minimal dynamic state to cope with changing content during training.
Across model sizes and data scales, RayDer showed strong power-law scaling and competitive zero-shot performance on many benchmarks.
The work points to a simpler, more scalable path for training view-synthesis models on large real-world video datasets without full supervision.
