Switch language한국어
Back to the list

Where to Look: Can Foundation Models Reach a Target Viewpoint Through Active Exploration?

TL;DR AI

Key summary

2 min read
  1. Researchers introduced Target Viewpoint Reproduction (TVR) and the TVRBench indoor benchmark to test whether foundation models can actively move in 3D space to match a target image.

  2. Current models perform poorly, especially when they must use multi-turn visual history or translate viewpoint gaps into movement rather than simple rotation.

  3. The study shows a clear weakness in active 3D spatial reasoning for both open- and closed-source models, including a 9B model.

  4. A unified post-training recipe combining visual-action supervised fine-tuning and multi-turn GRPO greatly improves success rates.

Read the original