CaMo: Camera Motion Grounded Evaluation and Training for Vision-Language Models

TL;DR AI
2 min readKey summary
A new paper finds that top vision-language models can ace direct spatial QA while still struggling with camera motion understanding.
The researchers propose the Spatial Narrative Score, which has models generate spatial narratives and then reason over them with a frozen proxy LLM.
They also introduce CaMo, a model trained to ground camera motion that performs more consistently across both benchmark styles.
The work suggests standard spatial benchmarks may overstate true 3D spatial intelligence and need better tests for transferable understanding.
