Meta's RankEvolve doubles down on agent reliability, lifting auto-research accuracy 45.8% to 62.5%
dair_ai · x · 2026-10-04
Meta released RankEvolve, a protocol for reliable auto-research agents: one silent bug (leaked eval data, disconnected gradient) can invalidate hours of training and every iteration built on it.
How it works:
- A compiled protocol enforces each research phase and gate
- Claude Code and Codex run as separate nodes that review and repair each other's changes
Results:
- At matched budget, combining both raises execution accuracy from 45.8% (best single product) to 62.5%
- Over 12 iterations on the open-source HSTU recommender, NDCG@10 on MovieLens-20M improved 4.48% over the published result
More from coding & agent
- Vercel hits $600M annualized revenue, up 148%, as coding agents drive half of new business — evilrabbit_ · 2026-10-04
- Hallmark: an open-source skill making Claude Code, Cursor and Codex UIs look less AI-generated — tom_doerr · 2026-10-04
- Stop fine-tuning to fix retrieval problems: Oracle technologist on where knowledge should live — AI Engineer · 2026-10-04
- KMP: recovering project decisions and evidence across Claude and Codex via MCP — Mountain-Raise-4556 · 2026-10-04
- (Lean)DOOM: DOOM fully rewritten in the Lean proof assistant, with formal proofs included — akbirthko · 2026-10-04
- Jin: a minimalist coding agent that swaps MCP/plugins for prompts and bash — aldapsiger · 2026-10-04