Switch language한국어
Back to the list

GPT-5.5 tops benchmarks but still hallucinates frequently at a 20 percent higher API cost

TL;DR AI

Key summary

2 min read
  1. OpenAI’s GPT-5.5 topped Artificial Analysis’ intelligence index, beating Claude Opus 4.7 and Gemini 3.1 Pro Preview.

  2. It used fewer tokens than GPT-5.4, but its effective API cost still rose by about 20%.

  3. Despite stronger benchmark results, separate testing found GPT-5.5 hallucinates frequently and performs poorly on BullshitBench.

  4. The takeaway: better scores and token efficiency do not necessarily mean the model is reliable in real-world use.

Read the original