Meta-reasoning harness hits 71.5% on ProgramBench, beating Codex by 13.5 points
anirudhg9119 · x · 2026-10-01
A new arXiv paper, Thinking Before Thinking: Scaling Agentic Inference Through Meta-Reasoning, from Paras Dahal, Anirudh Goyal and colleagues (including Taco Cohen, Rob Fergus, Jason Weston, Ruslan Salakhutdinov) proposes agentic meta-reasoning — an inference-time harness that makes execution control an explicit reasoning process.
- A controller consolidates what the run has established, explores next options, values each under the remaining budget, and dispatches work with context from persistent memory; workers do task-level computation. Between decisions the controller keeps only a compact run account, not full history.
- On ProgramBench (long-horizon agent capability via program reconstruction), meta-reasoning scores 71.5% with GPT-5.5 vs 58.0% for Codex, and 67.2% with Opus 4.8 vs 65.5% for Claude Code. Other benchmarks (abstract reasoning, multi-domain long-horizon reasoning, proof generation) show gains of 3.6–4.2 points over direct control.
- The framing: scaling reasoning increasingly means scaling the ability to structure computation itself — what to run, how it should depend on prior results, and when to stop.
Related event: Meta Proposes Agentic Meta-Reasoning for Self-Organized Inference Graphs(3 posts)→
More from coding & agent
- Merge Agent Handler ships as a partner recipe in NVIDIA NemoClaw for safe enterprise agents — shensi · 2026-10-01
- Google's Data Agent Kit hits GA, wiring 15+ data services into your coding agent via MCP — rseroter · 2026-10-01
- Claude Code 2.1.286 ships 88 CLI changes including secret-leak fix — ClaudeCodeLog · 2026-10-01
- ClawCast to Stream Episode on Jev Decision Model and OpenClaw Enterprise Agents — heyneighbor · 2026-10-01
- dcl-agent-mcp: open-source MCP server gives AI agents a real avatar in Decentraland — art2meta · 2026-10-01
- Open-sourced song-ad skill turns Claude into a 7-10 minute drama ad factory — SucceededMind · 2026-10-01