VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System
TL;DR AI
2 min readKey summary
VideoCoCo is a new text-to-video system that turns prompts into executable Blender code as an intermediate reasoning step.
A simulator produces a deterministic draft video, and a generative editor refines it into a photorealistic final output.
The team also released VideoCoCo-3K to train the editor on draft-instruction-target pairs and reported gains over a baseline on PhyGenBench and VBench-2.0.
The approach improves temporal and physical consistency by making scene dynamics explicit, executable, and easier to inspect.
