LLMs crush coding and math but choke on casual questions, and that's not a contradiction

TL;DR AI
2 min readKey summary
Latest models like GPT-5.4 Thinking and Claude Opus 4.6 are being used with Codex and Claude Code for work.
They can solve complex programming tasks in hours, restructure codebases, and even find vulnerabilities on their own.
But the same models can still fail on basic everyday questions, including in Advanced Voice Mode.
Andrej Karpathy says that is not a contradiction: code and math have clear checks and verifiable rewards.
Writing and consulting are harder to optimize because they lack clean metrics; tasks become easier to automate as they become more verifiable.



