Introducing SimpleQA

TL;DR AI
2 min readKey summary
OpenAI introduced SimpleQA, a 4,326-question benchmark for factuality evaluation.
The dataset uses short, fact-seeking questions with independently verified reference answers and easy grading.
It was filtered for single, stable answers and cross-checked by multiple trainers, with an estimated 3% error rate.
SimpleQA is meant to be less ambiguous than older benchmarks like TriviaQA and NQ, and more challenging for frontier models such as GPT-4o and o1.



