DeepSWE, a benchmark that prevents cheating by coding AI and enables more accurate measurement of programming performance

TL;DR AI
2 min readKey summary
Datacurve introduced DeepSWE, a new coding benchmark built from original tasks across 91 active open-source repositories and five programming languages.
The benchmark is designed to reduce training-data leakage and common loopholes, aiming for a fairer test of coding agents.
Datacurve says DeepSWE has lower validation error rates than SWE-Bench Pro and is harder to game.
In its tests, GPT-5.5 ranked highest, with the benchmark also referencing models from OpenAI, Poolside, and Claude.
