Switch language한국어
Back to the list

OpenAI successfully improves GPT-5.6 harness to triple ARC-AGI-3 score, showing that harnesses matter as much as the model itself

TL;DR AI

Key summary

2 min read
  1. OpenAI said GPT-5.6 Sol underperformed on ARC-AGI-3 because the evaluation harness was dropping its reasoning and truncating context too aggressively.

  2. With reasoning retention and compression instead of rolling truncation, the score jumped from 13.3% to 38.3%, while output tokens fell to one-sixth.

  3. The case shows benchmark results can depend heavily on API settings and harness design, not just model capability.

  4. That means AI performance comparisons need to account for how the evaluation is run, not only which model is tested.

Read the original