Even the latest AI models make three systematic reasoning errors, ARC-AGI-3 analysis shows

TL;DR AI
2 min readKey summary
ARC Prize Foundation analyzed 160 replays and traces from GPT-5.5 and Opus 4.7 on ARC-AGI-3.
They found three recurring failures: models miss the full world model, over-map tasks to familiar games, and solve levels without understanding why.
The results show frontier models still struggle with robust general reasoning in novel interactive environments.
That helps explain why ARC-AGI-3 scores remain below 1% and where current systems break down.
