SWE-Race: 188 real concurrency bugs benchmark where GPT-5.6 Luna scores 81%
heyitsdannyle · reddit · 2026-10-06
A new coding-agent benchmark, SWE-Race, is built from 188 real concurrency bugs (races, deadlocks, cancellations) merged across 100 Python projects. Each task is graded by the project's own tests in a network-free container with the repo cut to a single commit to prevent recovering fixes from git history. With one attempt per task, GLM-5.3 Flash scores 85% vs GPT-5.6 Luna's 81%; half the tasks are near-solved by all models, and the hard half shows 50%/45%/23% splits. The authors audited all 11k agent commands — 69 network attempts, all blocked; GLM tried 50 times to pip install the fixed release. Contamination checks (pre/post-2026 bugs) show a 9-point gap with confidence interval crossing zero. Half the tasks are private and public/private scores align. Dataset on Hugging Face, protocol follows DeepSWE's 100-step limit.
More from coding & agent
- A Reddit agent postmortem: the ERP's note-wiping behavior dictated the guardrails — max_gladysh · 2026-10-06
- Replit CEO: someone left an AI agent running overnight and woke up to $10,000 of wasted tokens — amasad · 2026-10-06
- Replit's Shlomi Fruchter: MCP, Harness and Skills are absurd concepts doomed like prompt engineering — shlomifruchter · 2026-10-06
- One Prompt Dump, Full App: Developer Wowed by Spawn Agent's Single-Shot Build — TAbrodi · 2026-10-06
- After Five Years, MATHPETS Launches as a Language and IDE for Agent-Based Models — jessi_cata · 2026-10-06
- Reading the source of 7 LLM eval tools uncovered 13 scoring bugs, 6 fixes merged — maverick_man1111 · 2026-10-06