SWE-sweep: new benchmark tasks agents with finding bugs before users hit them
klieret · reddit · 2026-10-03
Researchers from Meta, Stanford, Harvard and UW released SWE-sweep, a benchmark testing whether models can find and fix bugs before any user encounters them.
- Setup: hand an agent a full codebase and ask it to find & fix as many bugs as possible; scoring is based on a hidden set of known real-world bugs in the repos.
- Rigorous filtering ensures bugs are actually discoverable and fixable from reading the code alone.
- Early observations: scores are much lower than expected; fixing all of numpy is legitimately superhuman, but models also underperform on many small repos. Luna xhigh currently leads on cost efficiency.
- Fully open source (MIT), with more open-weight models to be added to the leaderboard soon.
More from coding & agent
- Fable 5.5 hallucinates less and replicates small models like Jev in hours, claims Bindu Reddy — bindureddy · 2026-10-03
- Anthropic engineer: future models will get much better at code deletion and simplification — simpsoka · 2026-10-03
- The Best AI Workflows Keep Friction Exactly Where Mistakes Matter — alifcoder · 2026-10-03
- Dev builds browser 3D game from scratch with Opus 5.5, Blender and Three.js — jason_mayes · 2026-10-03
- Cloudflare Durable Objects now survive client disconnects for long-running agents — threepointone · 2026-10-03
- Dev tired of nbviewer crashing builds serverless browser Jupyter notebook renderer, MIT-licensed — cneuralnetwork · 2026-10-03