Switch language한국어
Back to the list

Think, Act, Build: An Agentic Framework with Vision Language Models for Zero-Shot 3D Visual Grounding

TL;DR AI

Key summary

2 min read
  1. The paper presents Think, Act, Build (TAB), an agentic framework that uses 2D VLMs plus multi-view geometry for 3D visual grounding.

  2. TAB operates on raw RGB-D streams and reformulates 3D-VG as a 2D-to-3D reconstruction task driven by a VLM agent.

  3. The authors introduce Semantic-Anchored Geometric Expansion to propagate target locations across frames and aggregate multi-view features into 3D coordinates.

  4. Evaluations on ScanRefer and Nr3D show TAB outperforms prior zero-shot methods and exceeds some supervised baselines.

Read the original