Why Far Looks Up: Probing Spatial Representation in Vision-Language Models

TL;DR AI
2 min readKey summary
Researchers found that vision-language models often entangle vertical image position with perceived distance, revealing a persistent spatial bias across model families.
Using minimal contrastive pairs, the team tested how spatial information is represented internally rather than relying only on benchmark scores.
The internal vertical-distance entanglement predicted drops in accuracy and robustness, even when models looked strong on standard tasks.
To surface the problem more clearly, they introduced SpatialTunnel, a synthetic benchmark that removes common correlations from natural images.
