SWE-sweep: 4,100 real bugs across 100 repos, best model fixes only 4.7% at $7,230
OfirPress · x · 2026-10-02
Researchers from Meta Superintelligence Labs, Harvard, UW and Stanford launched SWE-sweep, a benchmark where agents must autonomously discover and fix as many bugs as they can in real repositories — with no hints about bug type or location. It covers 100 repositories and 4,100 bugs.
Highlights from the first leaderboard (all running mini-SWE-agent):
- Sol 5.6 (xhigh) tops the board at 4.7% bugs resolved, but costs $7,230 total ($72.30 per repo, 233 turns)
- Luna 5.6 (xhigh) resolves 2.5% for just $224 — the best cost-efficiency
- Opus 5 (xhigh) resolves 1.3% at $5,363; Kimi K3 resolves 0.6% at $2,451
The code is fully open source, and the team says automatic checks guard against fixes introducing regressions. The benchmark shows frontier models are still very weak at autonomously hunting bugs in large real-world codebases.
Related event: SWE-sweep benchmark: top models fix under 5% of bugs autonomously(3 posts)→
More from coding & agent
- Three reasons vibe-coded software is still far from production grade, with ReactBench data — aidenybai · 2026-10-02
- GitHub Copilot adds computer use to control desktop apps in public preview — PaulShellDev · 2026-10-02
- Dev builds her first Game Boy-style game entirely with OpenAI agents — craigsdennis · 2026-10-02
- When the Trace Looks Fine but the Agent Output Is Wrong — Sensitive-Parsnip-12 · 2026-10-02
- iPad + Tailscale + Codex is this dev's new favorite way to work outside the office — flavioAd · 2026-10-02
- Will Cloud AI Agents Need Residential IPs? Amazon Blocked Meta's Muse — SignificantFail3632 · 2026-10-02