GPT-5.5 tops benchmarks but still hallucinates frequently at a 20 percent higher API cost

TL;DR AI
2 min readKey summary
OpenAI’s GPT-5.5 topped Artificial Analysis’ intelligence index, beating Claude Opus 4.7 and Gemini 3.1 Pro Preview.
It used fewer tokens than GPT-5.4, but its effective API cost still rose by about 20%.
Despite stronger benchmark results, separate testing found GPT-5.5 hallucinates frequently and performs poorly on BullshitBench.
The takeaway: better scores and token efficiency do not necessarily mean the model is reliable in real-world use.



