SWE Refactor Bench: Coding Agents Pass Only 5.4% of Repo-Wide Migrations

The new SWE Refactor Bench tests coding agents on 20 real repo-wide stack migration tasks; across 520 runs only 5.4% passed, with the best model Opus 5 scoring just 47, exposing major gaps in long-horizon refactoring.

2026-08-25 ~ 2026-08-26 · 2 related posts