SWE Refactor Bench: rewriting 867K lines of C to Rust, Claude verifier rejects over 50% of agent submissions
jiqizhixin · x · 2026-09-04
Einsia AI's Navers Lab released SWE Refactor Bench, targeting long-horizon whole-repository stack migration that existing coding benchmarks miss.
- Scale: 20 real open-source projects, 867,000 lines of code, 10,594 files — rewrite entire systems from C to Rust with identical interfaces, behavior, and full deletion of old code.
- Industrial-grade repos: includes SQLite, zlib, libsodium, and GraphHopper.
- Verification pipeline: agents must pass migration review, behavioral equivalence testing, and adversarial vulnerability checking, not just compile.
- Early results: the strongest verifier (Claude Opus-5) breaks over 50% of agent submissions; other models only manage about 24%. The median time for a verifier to find a vulnerability was cut off in the source.
More from coding & agent
- Live show to cover rogue agent swarm incidents, GPT-6 Astra, and Runway's Solaris world model — DhruvBatra_ · 2026-09-05
- One Weekend Exercise for Learning to Build AI Products: Automate a Workflow End to End — realmadhuguru · 2026-09-05
- After Datadog and Grafana, dev endorses Pydantic Logfire for all observability — samuelcolvin · 2026-09-05
- Vibe coding isn't the problem—conflating it with agentic engineering is — bendee983 · 2026-09-05
- Allie Miller shares her AI research workflow: hypothesis-first with hundreds of agents — alliekmiller · 2026-09-05
- Builder shares update on Grok bot + Shopify integration experiment — billyjhowell · 2026-09-05