Switch language한국어
Back to the list

OpenAI Says Two API Settings Tripled GPT-5.6 Sol's ARC-AGI-3 Score

TL;DR AI

Key summary

2 min read
  1. OpenAI said two Responses API settings, retained reasoning and compaction, lifted GPT-5.6 Sol’s ARC-AGI-3 public-set score from 13.3% to 38.3%.

  2. The company also said the changes cut output-token usage by roughly 6x.

  3. OpenAI blamed the original benchmark harness for stopping the model from carrying useful reasoning across interactive puzzle tasks.

  4. The result highlights how agent benchmark scores can depend on API settings and harness design, not just model weights.

Read the original