Seeing Isn't Knowing: Do VLMs Know When Not to Answer Spatial Questions (and Why)?
TL;DR AI
2 min readKey summary
Researchers introduced SpatialUncertain, a benchmark for testing uncertainty handling in vision-language models on spatial reasoning tasks.
Across frontier open- and closed-source models, accuracy dropped sharply under occlusion and perspective ambiguity.
Models often failed to abstain when evidence was insufficient and were poor at identifying useful additional viewpoints.
The findings show vision-language models can be confidently wrong on spatial questions, raising safety concerns for real-world use.
