Why SWE-bench Verified no longer measures frontier coding capabilities | Hacker News

TL;DR AI
2 min readKey summary
Hacker News commenters argued SWE-bench Verified no longer reflects frontier coding ability.
They pointed to data contamination, benchmark-specific tuning, and inconsistent model claims as evidence of benchmark gaming.
The thread suggested that once public benchmarks are widely optimized against, they become weaker measures of real-world coding skill.
As an alternative, commenters proposed novel, periodically refreshed evaluation formats that are harder to overfit, similar to newer benchmark efforts like ARC-AGI.



