OpenAI claims GPT-5.6 Sol beats Opus 5 on ARC-AGI-3 with its latest API and two additional settings

TL;DR AI
2 min readKey summary
OpenAI said GPT-5.6 Sol scored 38.3% on ARC-AGI-3 using its Responses API with Retained Reasoning and Compaction, topping Claude Opus 5’s 30.2%.
The official benchmark harness gave a much lower 7.8% result, showing how test settings can dramatically change outcomes.
ARC Prize and François Chollet said provider-specific settings are acceptable if clearly disclosed, but standardized testing is still important for fair comparison.
The dispute underscores growing concerns about benchmark parity and how API features can affect cross-model comparisons.


