Switch language한국어
Back to the list

GPT-5.5 tops benchmarks but still hallucinates frequently and costs 20 percent more over the API

TL;DR AI

Key summary

2 min read
  1. OpenAI’s GPT-5.5 now tops the Artificial Analysis rankings, beating rivals like Claude Opus 4.7 and Gemini 3.1 Pro Preview on major benchmarks.

  2. It uses fewer tokens than GPT-5.4, but the doubled list price still makes its effective API cost about 20% higher.

  3. Despite the stronger scores, evaluations say GPT-5.5 still hallucinates frequently.

  4. It also did poorly on BullshitBench, where it often accepted or rationalized nonsensical prompts instead of rejecting them.

Read the original