SWE-sweep benchmark: top models fix under 5% of bugs autonomously

Meta, Harvard, UW and Stanford released SWE-sweep, a benchmark where agents must autonomously find and fix real bugs across 100 repositories without hints; the best model fixes only 4.7%. The code is open-sourced on GitHub.

2026-10-01 ~ 2026-10-02 · 3 related posts