Switch language한국어
Back to the list

Scale Labs debuts new Refactoring Leaderboard for AI

TL;DR AI

Key summary

2 min read
  1. Scale Labs launched the Refactoring Leaderboard as the final part of SWE Atlas to test AI coding agents on real-world refactoring in large codebases.

  2. The benchmark measures behavior-preserving edits, cleanup, and maintainability across multiple files, not just narrow coding prompts.

  3. Claude Code with Opus 4.7 currently ranks first, with ChatGPT 5.5 in second place.

  4. The research finds frontier closed models outperform open models, while reliability remains a major weakness because results can vary across repeated runs.

  5. The new benchmark raises the bar for enterprise use by focusing on code quality, repeatability, and production-style engineering work.

Read the original