Meta's RankEvolve: multi-agent cross-checking lifts auto-research execution accuracy to 62.5%
_reachsumit · x · 2026-10-01
Meta's RankEvolve (arXiv:2609.39551) is an auto-research framework for evolving generative ranking models. An Executable Operating Protocol compiles phases, gates, and loops into a runtime-enforced state machine, while a meta-meta-harness composes black-box coding agents (Claude Code, Codex) as execution-graph nodes that review and repair each other's work. Budget-matched results: heterogeneous composition raises all-oracle execution accuracy from 45.8% (best single product) to 62.5% (+16.7 points, 95% CI [6.6, 26.7]) with a 10.4% silent critical-defect rate; a knowledge layer carries findings including negative results across iterations. In a 12-iteration deployment on the open-source HSTU recommender it hit NDCG@10 of 0.2192 on MovieLens-20M LARGE (+4.48% over the published anchor) and 0.1948 on BASE (+2.80%).
More from coding & agent
- NVIDIA's Mid-Harness Scales Actions at the Model-Harness Boundary, Lifting TerminalBench Pass@1 to 68.03% — nvidia · 2026-10-01
- Amazon's SMART Self-Evolving Multi-Agent System Tops All 15 Subtitle Arena Directions, Cuts Penalty 6.9% — amazon · 2026-10-01
- Gary Bernhardt hits all-time low faith in AI agents: they "fix" tests by deleting them — sidjustice_ · 2026-10-01
- 'Agents are the software now' — developer urge to dive in — PurzBeats · 2026-10-01
- Hybris MCP Server lets AI assistants manage SAP Commerce Cloud instances — modelcontextprotocol · 2026-10-01
- cua-speedrun: CMU benchmark shows 4.4x speed gap between equal-scoring computer-use agents — arankomatsuzaki · 2026-10-01