Meta paper: dual coding agents cross-review lift correct patches from 45.8% to 62.5%
rohanpaul_ai · x · 2026-10-06
A new Meta paper finds that having two different coding agents review each other's patches catches far more silent bugs than giving one agent a bigger budget.
- Mixing Claude Code and Codex on the same task raised fully correct patches from 45.8% to 62.5% at matched spend.
- Agents editing real training code can leak test data, break a gradient, or miswire a flag — the code still runs, so GPU hours get burned on results that look valid but aren't.
- The mechanism is uncorrelated mistakes: agents from the same product fail the same way, leaving little for review to catch.
- The effect replicated on an unrelated training codebase.
Paper: "RankEvolve: A Reliable Multi-Agent Auto-Research Harness for Evolving Ranking Models" (arxiv. org/abs/2609.39551)
More from coding & agent
- HyperBrowseComp: a 423-question, 13-language stress test for web-browsing agents — zmkzmkz · 2026-10-06
- Cloudflare launches cf, an agent-first CLI to query observability data via the API — dinasaur_404 · 2026-10-06
- celld: Self-Hosted Cloudflare Workers-Style Runtime With a Cell per AI Agent — letandrewcook · 2026-10-06
- 9 Prompt Rules Cut Agent Thinking Up to 29% With Zero Task Loss, Across 664 Runs — PilgrimofHaqq2 · 2026-10-06
- Low Effort Can Burn More Tokens: Qwen3.8-27B-pi Fine-Tunes Coding Agent Effort Ordering — lmoroney · 2026-10-06
- Building a local LLM agent stack on a 128GB Mac Studio: Reddit thread weighs inference layer options — DrainBramage · 2026-10-06