VLMs are Good Teachers for Video Reasoning via Adaptive Test-Time Optimization
TL;DR AI
2 min readKey summary
Researchers proposed a test-time optimization framework that uses a vision-language model to infer task rules and convert them into differentiable rewards.
Those rewards guide a video generation model by updating a lightweight LoRA module during inference, without changing the base model.
On VBVR-Bench and RULER-Bench, the method improved average performance by 16.7 points and beat prior VLM-as-solver and Best-of-N approaches.
The result suggests VLMs can act as effective teachers and evaluators for stronger video reasoning at test time.
