Why Far Looks Up: Probing Spatial Representation in Vision-Language Models
TL;DR AI
2 min readKey summary
Researchers found a persistent spatial bias in vision-language models: vertical image position is often tied to perceived distance.
The bias resembles perspective cues in natural photos, but it creates gaps between normal and counter-heuristic cases.
As models scale and benchmark scores improve, this shortcut can persist or even strengthen, masking weak spatial reasoning.
To reveal it, the team introduced SpatialTunnel, a synthetic benchmark that strips away common image correlations.
Models with more disentangled spatial axes were more reliable across spatial benchmarks and showed better robustness.
