See2Think: Do Multimodal Models Really Use Intermediate Visual States?
TL;DR AI
2 min readKey summary
Researchers introduced See2Think, a 1,200-problem benchmark paired with a Visual Action-of-Thought protocol.
The framework tests whether multimodal LLMs truly rely on intermediate images such as sketches and annotations during reasoning.
Results showed strong variation across models and settings, with faithful visual rendering emerging as a major weakness.
Corrupted visual feedback can cut accuracy by more than 10 points, showing how sensitive visual reasoning is to intermediate state quality.
