DeepSWE: A Contamination-Resistant Coding Benchmark — James Shi, Datacurve
DeepSWE is a contamination-resistant coding benchmark with 113 original tasks across ~100 repositories
Datacurve's DeepSWE benchmark addresses contamination and saturation problems in existing coding benchmarks like SWEBench Pro by using 113 hand-crafted original tasks spread across nearly 100 repositories, making it harder for models to cheat via training data memorization. It has already replaced SWEBench Pro in the Artificial Analysis coding agent index and is being used by frontier model labs. The benchmark spans TypeScript, JavaScript, Python, Rust, and Go with plans for more languages.