Switch language한국어
Back to the list

PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models

TL;DR AI

Key summary

2 min read
  1. Researchers introduced PerceptionBench, a benchmark built from 42 existing benchmarks to test 10 atomic visual perception skills in multimodal large language models.

  2. The dataset includes 3,000 verified short-answer questions designed to separate perception failures from reasoning and knowledge errors.

  3. Testing 16 frontier MLLMs found that none exceeded 60% accuracy, and perception-related hallucination was the weakest area on average.

  4. The results show that similar overall scores can hide major differences in capability profiles, making diagnosis of visual perception failures more precise.

Read the original