Incantation: Natural Language as the Action Interface for Multi-Entity Video World Models

TL;DR AI
2 min readKey summary
Researchers introduced Incantation, an interactive video world model that uses natural language every 0.25 seconds for fine-grained control.
It can control multiple entities at once and shows stronger cross-entity transfer and prompt robustness than an action-index baseline.
The paper reports that a lightweight student model can run in real time, helping support faster long-horizon generation.
The approach points to a more flexible action interface for interactive video models across different systems and tasks.
