Switch language한국어
Back to the list

How and What to Imagine? Visual Thinking in Unified Multimodal Models for Cross-View Spatial Reasoning

TL;DR AI

Key summary

2 min read
  1. Researchers propose View Dropout, which hides parts of one view from the answer path while keeping them visible for visual thinking tokens.

  2. They compare top-down, panoramic, and point-matching visual reasoning, and find panoramic thinking is the easiest to learn and most useful for reasoning.

  3. View Dropout plus panoramic visual thinking is the only setup that consistently improves cross-view spatial reasoning.

  4. The approach delivers the best results on five out-of-domain benchmarks, showing stronger generalization beyond training conditions.

Read the original