SWE Refactor Bench: Only 5.4% of 520 Runs Pass, Best Model Opus 5 Scores 47/100

YouJiacheng · x · 2026-08-26

A new benchmark, SWE Refactor Bench (arXiv:2608.23564), tests whether coding agents can autonomously complete long-horizon, whole-repository migrations to pay down technical debt.

Existing benchmarks only check behavioral correctness, enabling an easy hack—copying the original implementation to pass tests—termed Blindness by the authors. The benchmark includes 20 whole-repo migrations across 4 kinds of technical debt, with a three-stage protocol: (1) Migration Audit verifies the migration actually occurred; (2) Behavioral Tests measure correctness with a fixed suite; (3) Agentic Verification uses 6 independent coding agents to generate targeted tests for hidden behavioral differences.

Across 520 runs from 8 frontier models and 26 model-effort configurations, only 28 (5.4%) pass all three stages; 13 of 20 tasks receive no accepted solution; the best model, claude-opus-5, scores just 47.0/100. The paper argues migration completeness and behavioral correctness are distinct abilities.

Related event: SWE Refactor Bench: Coding Agents Pass Only 5.4% of Repo-Wide Migrations(2 posts)→

Original post →

More from coding & agent

coding & agent channel →