Meta paper: branching self-improving harness search lifts Olympiad math accuracy to 62%
rohanpaul_ai · x · 2026-10-04
A new paper from Meta, Duke and University of California introduces a branched approach to self-improving agent harness optimization:
- A harness is the code around an LLM controlling tools, retrieval and self-checks. Prior work (Meta-Harness) rewrites it iteratively against a fixed dev set, so search converges to one path and risks local optima.
- The new method splits search into branches, each keeping the dev cases it solves better than others and writing its own notes on what worked — in math, one branch learned answer verification while another learned full derivations.
- A router selects the best-fit branch head per new input before execution.
- Results: on Gemini 3 Flash, Olympiad-level math accuracy rose from 46.0% to 62.0% (34.8% relative gain), plus +11.6% on Terminal-Bench 2.0 and +3.8% on SWE-bench Lite, beating Meta-Harness in all 4 settings.
Practical takeaway: when auto-tuning agents, keep several specialized harnesses and route between them rather than betting on a single winner.
More from coding & agent
- Meta's RankEvolve doubles down on agent reliability, lifting auto-research accuracy 45.8% to 62.5% — dair_ai · 2026-10-04
- CMU paper: RL-trained 4B proposer edits agent harness code, beats its 35B teacher — omarsar0 · 2026-10-04
- Dev praises pi-durable for running long-lived agents behind any interface — tokumin · 2026-10-04
- Claude Code 2.1.289 prompt tokens jump 17.8%, system prompts now over half — ClaudeCodeLog · 2026-10-04
- Claude Code 2.1.289 patches @-mention symlink bypass of file read deny rules — ClaudeCodeLog · 2026-10-04
- Claude Code 2.1.289 ships agent.spawn and fixes @-mention permission bypass — ClaudeCodeLog · 2026-10-04