SWE-sweep: 4,100 real bugs across 100 repos, best model fixes only 4.7% at $7,230

OfirPress · x · 2026-10-02

Researchers from Meta Superintelligence Labs, Harvard, UW and Stanford launched SWE-sweep, a benchmark where agents must autonomously discover and fix as many bugs as they can in real repositories — with no hints about bug type or location. It covers 100 repositories and 4,100 bugs.

Highlights from the first leaderboard (all running mini-SWE-agent):

The code is fully open source, and the team says automatic checks guard against fixes introducing regressions. The benchmark shows frontier models are still very weak at autonomously hunting bugs in large real-world codebases.

Related event: SWE-sweep benchmark: top models fix under 5% of bugs autonomously(3 posts)→

Original post →

More from coding & agent

coding & agent channel →