SWE Refactor Bench reveals 5.4% survival rate for coding agents in whole-repo migrations
_akhaliq · x · 2026-08-25
Einsia released SWE Refactor Bench, a benchmark for long-horizon, whole-repository software stack migrations. It includes 20 real-world tasks like porting SQLite or zlib. Out of 520 runs, only 28 passed all stages, a 5.4% survival rate, with 13/20 tasks solved by no one. This indicates that while agents are good at bug fixes, system-scale refactoring remains an open challenge.
More from coding & agent
- LangChain's Harrison Chase ships a skill for building agent evals iteratively — hwchase17 · 2026-08-25
- LangChain open-sources eval-engineering skill: building agent eval environments from traces and human feedback — BraceSproul · 2026-08-25
- Why Has No One Built an Agent-Native Email Experience Yet? — maddiehfaulkner · 2026-08-25
- MacStories: M6 and M5 Ultra offer huge potential for local AI on macOS — Dimillian · 2026-08-25
- GPT-OSS 20B Ran My Personal Agent for a Week: 312 Tasks, 97.4% First-Shot Tool Calls, Zero Frontier APIs — rasheed106 · 2026-08-25
- Ask HN-Style: Is There a Debloated Local LLM Optimized Purely for Web Dev? — HsSekhon · 2026-08-25