Switch language한국어
Back to the list

Why Far Looks Up: Probing Spatial Representation in Vision-Language Models

TL;DR AI

Key summary

2 min read
  1. Researchers found that vision-language models often entangle vertical image position with perceived distance, revealing a persistent spatial bias across model families.

  2. Using minimal contrastive pairs, the team tested how spatial information is represented internally rather than relying only on benchmark scores.

  3. The internal vertical-distance entanglement predicted drops in accuracy and robustness, even when models looked strong on standard tasks.

  4. To surface the problem more clearly, they introduced SpatialTunnel, a synthetic benchmark that removes common correlations from natural images.

Read the original