New SWE-Refactor-Bench exposes AI coding weakness: 0% success in full repo migration
机器之心 · wechat · 2026-08-27
Jiqizhixin reports on the release of a new benchmark, SWE-Refactor-Bench, which escalates testing from bug fixes to full repository refactoring (e.g., rewriting SQLite from C to Rust). It covers 20 real-world projects with 867k lines of code. The evaluation features three stages: migration audit, 130k functional tests, and a red teaming phase where other agents attack the code. Results show that across 520 attempts by 8 frontier models, the success rate was only 5.4%, with 13 tasks unsolved. This reveals a core flaw in current AI agents: the difficulty of balancing "completing the migration" with "preserving exact behavior," especially the final 0.1% of correctness that prevents system-level failures.
Related event: SWE Refactor Bench: Coding Agents Pass Only 5.4% of Full-Repo Migrations(4 posts)→
More from coding & agent
- Claude Opus 5.5 Sweeps CursorBench at 57.8% Max, Costs 40% Less Per Task Than Opus 5 — mattyp · 2026-09-23
- Claude Caught Emitting Harmful Requests: Secret Exfiltration and Hostile CLAUDE.md Injections — maksym_andr · 2026-09-23
- Claude reportedly emits harmful requests, exfiltrates secrets via hostile CLAUDE.md text — maksym_andr · 2026-09-23
- Anthropic ships Opus 5.5: up to 40% cheaper than Opus 5, pulling Codex users back to Claude — danshipper · 2026-09-23
- Simple Post MCP App in ChatGPT Makes Social Posting Effortless — haltakov · 2026-09-23
- eve launches Workflow tools: deterministic subagent orchestration with durable multi-step execution — cramforce · 2026-09-23