“Don’t Trust AI Scores”: OpenAI Reveals the Fatal Limitations of Benchmarks

TL;DR AI
2 min readKey summary
OpenAI says standard Q&A benchmarks alone can’t fully measure modern AI capability or safety.
The company released a joint playbook urging evaluators to inspect not just the model, but also the harness, token budget, retry count, and other run conditions.
Performance can vary sharply with budget and environment, and results can be distorted by reward hacking or data contamination.
The guidance is meant to make third-party comparisons and regulatory decisions more accurate for frontier and agentic AI.



