Anthropic's Opus 5 blows past Fable 5 and GPT-5.6 Sol on the benchmark designed to measure real intelligence

TL;DR AI
2 min readKey summary
Anthropic’s Claude Opus 5 set a new high score on ARC-AGI-3 at 30.2%, far ahead of GPT-5.6 Sol’s 7.8%.
The model solved five previously unsolved environments, including four at or above human level.
Researchers said the gains appear to reflect better reasoning and planning, not just benchmark-specific tuning.
The result underscores progress on hard, unfamiliar tasks and intensifies competition on general reasoning benchmarks.



