Switch language한국어
Back to the list

CaMo: Camera Motion Grounded Evaluation and Training for Vision-Language Models

TL;DR AI

Key summary

2 min read
  1. A new paper finds that top vision-language models can ace direct spatial QA while still struggling with camera motion understanding.

  2. The researchers propose the Spatial Narrative Score, which has models generate spatial narratives and then reason over them with a frozen proxy LLM.

  3. They also introduce CaMo, a model trained to ground camera motion that performs more consistently across both benchmark styles.

  4. The work suggests standard spatial benchmarks may overstate true 3D spatial intelligence and need better tests for transferable understanding.

Read the original