SWE-sweep benchmark: top models score under 5% finding bugs unsupervised
OfirPress · x · 2026-10-01
- Ofir Press's team released SWE-sweep, a benchmark testing whether LMs can discover and fix bugs without being told what went wrong—key for agents that proactively maintain repos.
- It spans 100 real repos, 22 languages, and 4k real bugs (numpy, PHP interpreter, Lean kernel).
- Top models score <5%, showing labs haven't started climbing this ability curve.
- The authors see it as a north-star benchmark, among their most challenging ever.
More from coding & agent
- nanoGPT speedrun: AI agent hits val loss 3.28 in 2726 steps, nearing human record of 2600 — zsakib_ · 2026-10-01
- One prompt, done: agent autonomously debugs kernel and GPU compatibility — MaziyarPanahi · 2026-10-01
- Convert images to text with an LLM: Markitdown workflow tip #5 — mdancho84 · 2026-10-01
- Microsoft launches Markitdown, a free Python library converting any document to Markdown — mdancho84 · 2026-10-01
- Why Experts Get More From AI Agents: Give Them Bigger Chunks — kieranklaassen · 2026-10-01
- Matz is merging 90 PRs an hour and CI runners can't keep up — kieranklaassen · 2026-10-01