Are VLMs Seeing or Just Saying? Uncovering the Illusion of Visual Re-examination

TL;DR AI
2 min readKey summary
Researchers introduced VisualSwap and VS-Bench to test whether vision-language models truly re-examine images after reflective prompts.
Across multiple models, swapping the image caused large accuracy drops, with reasoning-heavy models especially vulnerable.
The findings suggest many VLMs may simulate reinspection in text while failing to stay visually grounded.
Multi-turn user prompts improved visual attention, but self-generated reflection did not, raising reliability concerns for image-based tasks.
