Meta's new SWE-sweep eval makes coding agents find bugs in real repos with no hints
jack_w_rae · x · 2026-10-02
Jack Rae shared a new eval from Meta's Code team called SWE-sweep, which inverts the premise of today's coding benchmarks.
- Most benchmarks hand the agent a concrete task (repo + GitHub issue), meaning humans already did the hard part of noticing and describing the bug.
- SWE-sweep removes the ticket: the agent gets a real repo built from 100 open-source projects with 4k+ real bugs and one instruction — find and fix as many bugs as you can.
- It measures something new: exploring a large codebase, distinguishing bugs from intended behavior, prioritizing fixes, and debugging without regressions in one long, open-ended run.
- The author says leaderboard results are very interesting, but the quoted tweet is truncated and specific scores are not included.
More from coding & agent
- Where should agent tool-call state live: model context or app? — LowMixture1084 · 2026-10-02
- Forward Deployed Engineers Earning $1M+/Year Share Their Full Playbook for Fortune 500 AI Rollouts — vasuman · 2026-10-02
- Agent Pipeline Cuts YouTube Shorts Production from 4-5 Hours to 20 Minutes — hugobowne · 2026-10-02
- Agent-to-agent collaboration is now the norm: every coding task mixes Codex and Claude agents — andreisavu · 2026-10-02
- Liquid AI's decision model d1 lands on Vercel AI Gateway at $0.04/M input tokens — JosephJacks_ · 2026-10-02
- Dev discovers his game-testing agent adds two humans to check mixed-reality accessibility for kids and adults — nptacek · 2026-10-02