RefCaptioner: Multi-Reference Image-Grounded Video Captioning
Key summary
Researchers introduced RefCaptioner, a video captioning system that uses multiple reference images to ground specific visual phrases more faithfully.
The paper defines a new task, multi-reference image-grounded video captioning, focused on factual, phrase-level descriptions rather than generic captions.
RefCaptioner uses a two-stage training pipeline with supervised fine-tuning and reinforcement learning to improve reference selection, grounding accuracy, and distractor rejection.
The authors also release a large training corpus and MRVBench for evaluation, reporting strong results on open-source benchmarks and favorable human judgments.
This pushes video captioning beyond generic descriptions toward captions that are more factual, verifiable, and useful for source-faithful video generation.
