SceneGraphVLM: Dynamic Scene Graph Generation from Video with Vision-Language Models

TL;DR AI
2 min readKey summary
Researchers introduced SceneGraphVLM, a two-stage vision-language system for generating scene graphs from images and videos.
It combines compact VLMs, token-efficient graph serialization, and hallucination-aware reinforcement learning to improve graph quality.
For video, the model can also use prior-frame graph context, helping it track objects and relations more consistently over time.
The approach aims to make structured visual understanding faster and more practical for downstream computer vision systems.
