SWE-Race: 188 real concurrency bugs benchmark where GPT-5.6 Luna scores 81%

heyitsdannyle · reddit · 2026-10-06

A new coding-agent benchmark, SWE-Race, is built from 188 real concurrency bugs (races, deadlocks, cancellations) merged across 100 Python projects. Each task is graded by the project's own tests in a network-free container with the repo cut to a single commit to prevent recovering fixes from git history. With one attempt per task, GLM-5.3 Flash scores 85% vs GPT-5.6 Luna's 81%; half the tasks are near-solved by all models, and the hard half shows 50%/45%/23% splits. The authors audited all 11k agent commands — 69 network attempts, all blocked; GLM tried 50 times to pip install the fixed release. Contamination checks (pre/post-2026 bugs) show a 9-point gap with confidence interval crossing zero. Half the tasks are private and public/private scores align. Dataset on Hugging Face, protocol follows DeepSWE's 100-step limit.

Original post →

More from coding & agent

coding & agent channel →