Benchmarking Opus 5 on SlopCodeBench | Hacker News
TL;DR AI
2 min readKey summary
A Hacker News discussion outlined plans to benchmark several Anthropic models on selected SlopCodeBench problems.
The evaluation would compare Claude Opus 5, Opus 4.8, Sonnet 5, and Fable using an article, related repos, and staged tasks.
The tasks are organized into easy, medium, and hard levels to test performance more systematically.
The post highlights ongoing interest in structured benchmarks for judging real-world coding model capability and choosing the right model.


