SWE-sweep benchmark: top models fix under 5% of bugs autonomously
Meta, Harvard, UW and Stanford released SWE-sweep, a benchmark where agents must autonomously find and fix real bugs across 100 repositories without hints; the best model fixes only 4.7%. The code is open-sourced on GitHub.
2026-10-01 ~ 2026-10-02 · 3 related posts
- SWE-sweep benchmark: top models score under 5% finding bugs unsupervised — OfirPress · 2026-10-01
- SWE-sweep: 4,100 real bugs across 100 repos, best model fixes only 4.7% at $7,230 — OfirPress · 2026-10-02
- SWE-sweep benchmark code open-sourced under facebookresearch on GitHub — OfirPress · 2026-10-02