Wonder: Video World Model Done Better
TL;DR AI
2 min readKey summary
Researchers introduced Wonder, a real-time video world model that turns an image or conditional video into a playable scene.
Users can navigate scenes with camera motion, re-render video-conditioned content, and preserve coherent appearance, geometry, and motion.
The system uses dense coordinate-field camera conditioning, sparse attention memory, and improved distillation to better follow controls and keep diversity.
Wonder can generate minute-scale videos at 16 FPS, advancing interactive simulation and long-horizon video creation.
