PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models
TL;DR AI
2 min readKey summary
Researchers introduced PerceptionBench, a benchmark built from 42 existing benchmarks to test 10 atomic visual perception skills in multimodal large language models.
The dataset includes 3,000 verified short-answer questions designed to separate perception failures from reasoning and knowledge errors.
Testing 16 frontier MLLMs found that none exceeded 60% accuracy, and perception-related hallucination was the weakest area on average.
The results show that similar overall scores can hide major differences in capability profiles, making diagnosis of visual perception failures more precise.
