The Expense of Seeing: Attaining Trustworthy Multimodal Reasoning Within the Monolithic Paradigm
TL;DR AI
2 min readKey summary
Researchers say many vision-language models rely too much on language shortcuts and do not truly fuse image and text.
The paper introduces the Modality Translation Protocol to test semantic sufficiency rather than simple multimodal performance gains.
New metrics like Toll, Curse, and Fallacy of Seeing are designed to measure whether models actually use visual information.
The work suggests current benchmarks may overestimate visual understanding and should be reconsidered for future model design.
