The Hallway Track
Research Findings

DeepSWE: A Contamination-Resistant Coding Benchmark — James Shi, Datacurve

AI Engineer · Jul 26, 2026 · Research Findings

DeepSWE is a contamination-resistant coding benchmark with 113 original tasks across ~100 repositories

Datacurve's DeepSWE benchmark addresses contamination and saturation problems in existing coding benchmarks like SWEBench Pro by using 113 hand-crafted original tasks spread across nearly 100 repositories, making it harder for models to cheat via training data memorization. It has already replaced SWEBench Pro in the Artificial Analysis coding agent index and is being used by frontier model labs. The benchmark spans TypeScript, JavaScript, Python, Rust, and Go with plans for more languages.

benchmarks coding-agents swe-bench evaluation contamination

Watch / read the original source →