AI benchmarks systematically ignore how humans disagree, Google study finds

TL;DR AI
2 min readKey summary
A Google Research and RIT study finds typical AI benchmarks using three to five raters per example lose important disagreement information.
Researchers built a simulator calibrated on five datasets to test thousands of budget splits between examples and raters.
They find fewer than ten raters per example often fails to produce reproducible model comparisons, while about 1,000 total annotations can suffice if split correctly.
The optimal allocation depends on the evaluation metric: accuracy favors many examples with few raters, distribution-aware metrics need fewer examples with many raters.



