SWE-sweep benchmark tests whether AI agents can find bugs before users hit them

klieret · reddit · 2026-10-02

Researchers from Meta, Stanford, Harvard, and UW launched SWE-sweep, a new benchmark that flips the usual SWE setup: instead of fixing bugs users already hit, an agent is handed a large real-world codebase and asked to find and fix as many bugs as possible, scored against a hidden set of real bugs.

Related event: Meta and Partners Release SWE-sweep: Top Model Fixes Only 4.7% of Hidden Bugs(6 posts)→

Original post →

More from coding & agent

coding & agent channel →