RayDer: Scalable Self-Supervised Novel View Synthesis from Real-World Video
TL;DR AI
2 min readKey summary
RayDer is a unified feed-forward transformer for self-supervised novel view synthesis.
It combines camera estimation, scene reconstruction, and rendering in one model, with a minimal dynamic state for changing content.
The system trains stably on unconstrained real-world videos and scales predictably with more data and larger models.
It delivers competitive zero-shot results across many benchmarks, pointing to better 3D understanding without costly labels.
