Switch language한국어
Back to the list

Why Far Looks Up: Probing Spatial Representation in Vision-Language Models

TL;DR AI

Key summary

2 min read
  1. Researchers found a persistent spatial bias in vision-language models: vertical image position is often tied to perceived distance.

  2. The bias resembles perspective cues in natural photos, but it creates gaps between normal and counter-heuristic cases.

  3. As models scale and benchmark scores improve, this shortcut can persist or even strengthen, masking weak spatial reasoning.

  4. To reveal it, the team introduced SpatialTunnel, a synthetic benchmark that strips away common image correlations.

  5. Models with more disentangled spatial axes were more reliable across spatial benchmarks and showed better robustness.

Read the original