Switch language한국어
Back to the list

Introducing SimpleQA

TL;DR AI

Key summary

2 min read
  1. OpenAI introduced SimpleQA, a 4,326-question benchmark for factuality evaluation.

  2. The dataset uses short, fact-seeking questions with independently verified reference answers and easy grading.

  3. It was filtered for single, stable answers and cross-checked by multiple trainers, with an estimated 3% error rate.

  4. SimpleQA is meant to be less ambiguous than older benchmarks like TriviaQA and NQ, and more challenging for frontier models such as GPT-4o and o1.

Read the original