Switch language한국어
Back to the list

OpenAI Says AI Benchmark Scores Depend on Harnesses, Budgets, and Memory Design

TL;DR AI

Key summary

2 min read
  1. OpenAI released guidance saying benchmark scores can shift a lot based on the test harness, compute budget, tools, and memory/context setup.

  2. The company says evaluators should disclose these conditions and separate capability claims from controlled comparisons and robustness tests.

  3. It points to GPT-5.5 cyber-range results, where context compaction improved long-task performance, as evidence that evaluation design changes outcomes.

  4. The message: benchmark scores should be read as results under specific testing setups, not fixed measures of a model’s standalone ability.

Read the original