SWE-sweep benchmark code open-sourced under facebookresearch on GitHub
OfirPress · x · 2026-10-02
The SWE-sweep benchmark code is now open source on GitHub under facebookresearch. It is a thin wrapper around Harbor; agents autonomously find and fix bugs in real repos without hints. The repo includes evaluation tasks, separate verifier environments, and micro-average scoring, and currently pins a specific Harbor revision since Harbor 0.23 lacks the required hooks.
Related event: SWE-sweep benchmark: top models fix under 5% of bugs autonomously(3 posts)→
More from Research
- ARPA-H plans to cut clinical trials from 10+ years to under 4 with AI — Afinetheorem · 2026-10-02
- Tacit-TTS: transcript-free voice cloning 10x faster than IndexTTS2 — Jian Chen · 2026-10-02
- ByteDance Seed's RWTD lifts one-step SANA Sprint GenEval from 0.73 to 0.80 — ByteDance-Seed · 2026-10-02
- Amazon's position-selective self-distillation trains LLM judges that beat RL by 2-9 points — amazon · 2026-10-02
- mhctools: one Python wrapper for a dozen MHC immunoinformatics predictors — iskander · 2026-10-02
- How to cost AI-powered filters: roofline model puts 5k-review LLM filter floor at 6.6s on H100 — sh_reya · 2026-10-02