OpenAI Says AI Benchmark Scores Depend on Harnesses, Budgets, and Memory Design

TL;DR AI
2 min readKey summary
OpenAI released guidance saying benchmark scores can shift a lot based on the test harness, compute budget, tools, and memory/context setup.
The company says evaluators should disclose these conditions and separate capability claims from controlled comparisons and robustness tests.
It points to GPT-5.5 cyber-range results, where context compaction improved long-task performance, as evidence that evaluation design changes outcomes.
The message: benchmark scores should be read as results under specific testing setups, not fixed measures of a model’s standalone ability.
