SWE-sweep benchmark tests whether AI agents can find bugs before users hit them
klieret · reddit · 2026-10-02
Researchers from Meta, Stanford, Harvard, and UW launched SWE-sweep, a new benchmark that flips the usual SWE setup: instead of fixing bugs users already hit, an agent is handed a large real-world codebase and asked to find and fix as many bugs as possible, scored against a hidden set of real bugs.
- All bugs are real-world issues
- Standout finding: Luna xhigh is highly cost-efficient, scoring nearly half of Sol at a tiny fraction of the cost, outperforming the tested Anthropic models on value
- Full leaderboard and construction details at swesweep.com, with an accompanying paper
More from coding & agent
- Dev argues AI is great at assets and code but bad at designing game mechanics that feel good — rms80 · 2026-10-03
- Free one-day curriculum takes you from AI agent basics to MCP and agent security — ifioknkem · 2026-10-03
- Using System One models in Swift: fast, deterministic decisions via Apple Foundation Models — rxwei · 2026-10-03
- Run 50 AI Coding Agents in Parallel With One Global Rule for Background Tasks — Daniel_Farinax · 2026-10-03
- Why Linear's Agent Session beats Slack as a collaboration surface for agentic work — jeff_weinstein · 2026-10-03
- SWE-chat V2 ships 3.5x larger with agent skills and subagent trajectories — Diyi_Yang · 2026-10-03