OpenAI successfully improves GPT-5.6 harness to triple ARC-AGI-3 score, showing that harnesses matter as much as the model itself

TL;DR AI
2 min readKey summary
OpenAI said GPT-5.6 Sol underperformed on ARC-AGI-3 because the evaluation harness was dropping its reasoning and truncating context too aggressively.
With reasoning retention and compression instead of rolling truncation, the score jumped from 13.3% to 38.3%, while output tokens fell to one-sixth.
The case shows benchmark results can depend heavily on API settings and harness design, not just model capability.
That means AI performance comparisons need to account for how the evaluation is run, not only which model is tested.
