Where to Look: Can Foundation Models Reach a Target Viewpoint Through Active Exploration?
TL;DR AI
2 min readKey summary
Researchers introduced Target Viewpoint Reproduction (TVR) and the TVRBench indoor benchmark to test whether foundation models can actively move in 3D space to match a target image.
Current models perform poorly, especially when they must use multi-turn visual history or translate viewpoint gaps into movement rather than simple rotation.
The study shows a clear weakness in active 3D spatial reasoning for both open- and closed-source models, including a 9B model.
A unified post-training recipe combining visual-action supervised fine-tuning and multi-turn GRPO greatly improves success rates.
