GPT-5.5 tops benchmarks but still hallucinates frequently and costs 20 percent more over the API

TL;DR AI
2 min readKey summary
OpenAI’s GPT-5.5 now tops the Artificial Analysis rankings, beating rivals like Claude Opus 4.7 and Gemini 3.1 Pro Preview on major benchmarks.
It uses fewer tokens than GPT-5.4, but the doubled list price still makes its effective API cost about 20% higher.
Despite the stronger scores, evaluations say GPT-5.5 still hallucinates frequently.
It also did poorly on BullshitBench, where it often accepted or rationalized nonsensical prompts instead of rejecting them.



