Switch language한국어
Back to the list

Why SWE-bench Verified no longer measures frontier coding capabilities | Hacker News

TL;DR AI

Key summary

2 min read
  1. Hacker News commenters argued SWE-bench Verified no longer reflects frontier coding ability.

  2. They pointed to data contamination, benchmark-specific tuning, and inconsistent model claims as evidence of benchmark gaming.

  3. The thread suggested that once public benchmarks are widely optimized against, they become weaker measures of real-world coding skill.

  4. As an alternative, commenters proposed novel, periodically refreshed evaluation formats that are harder to overfit, similar to newer benchmark efforts like ARC-AGI.

Read the original