CrossView Suite: Harnessing Cross-view Spatial Intelligence of MLLMs with Dataset, Model and Benchmark

TL;DR AI
2 min readKey summary
Researchers introduced CrossView Suite, a three-part system for cross-view spatial intelligence in multimodal large language models.
It includes CrossViewSet, a 1.6M-sample cross-view instruction dataset, and CrossViewBench, a scene-disjoint benchmark for evaluation.
The CrossViewer model uses a three-stage pipeline to align and fuse multi-view object features for better spatial reasoning.
The work addresses a major weakness in MLLMs: recognizing the same objects and scenes reliably across different viewpoints.
