VLMs are Good Teachers for Video Reasoning via Adaptive Test-Time Optimization

TL;DR AI
2 min readKey summary
Researchers propose using vision-language models as test-time teachers, not direct solvers, for video reasoning.
The method extracts task rules and progress signals from VLMs, then turns them into differentiable rewards.
These rewards guide online updates to a lightweight LoRA module in a video generation model.
The approach delivers strong gains over prior VLM-as-solver and Best-of-N methods on symbolic and general benchmarks.
