OpenAI Says Two API Settings Tripled GPT-5.6 Sol's ARC-AGI-3 Score

TL;DR AI
2 min readKey summary
OpenAI said two Responses API settings, retained reasoning and compaction, lifted GPT-5.6 Sol’s ARC-AGI-3 public-set score from 13.3% to 38.3%.
The company also said the changes cut output-token usage by roughly 6x.
OpenAI blamed the original benchmark harness for stopping the model from carrying useful reasoning across interactive puzzle tasks.
The result highlights how agent benchmark scores can depend on API settings and harness design, not just model weights.
