Meta-reasoning harness hits 71.5% on ProgramBench with GPT-5.5, beating Codex's 58.0%
rohanpaul_ai · x · 2026-10-01
The arXiv paper "Thinking Before Thinking: Scaling Agentic Inference Through Meta-Reasoning" (Paras Dahal et al., 12 authors) introduces agentic meta-reasoning: an inference-time harness that turns execution control into an explicit, structured reasoning process.
Design
- Workers perform task-level computation; a controller consolidates what the run has established, explores options, values each under the remaining budget, and dispatches work with context from persistent memory.
- Between decisions the controller carries only a compact account of the run instead of replaying full history.
- Targets the control problem in long, complex agent runs: which partial work to build on, whether to restart, when to stop.
Results
- Baselines include production coding agents, research harnesses, and a Direct Control Agent with the same workers and compute budget.
- On ProgramBench (long-horizon capability via program reconstruction): 71.5% with GPT-5.5 vs 58.0% for Codex; 67.2% with Opus 4.8 vs 65.5% for Claude Code.
- Gains of 3.6–4.2 points over direct control on benchmarks spanning abstract reasoning, multi-domain long-horizon reasoning, and proof generation.
Related event: Meta Proposes Agentic Meta-Reasoning Framework for Scaling Inference(4 posts)→
More from coding & agent
- The real agent product is knowing when to stop, not just acting autonomously — SucceededMind · 2026-10-01
- LiteLLM launches Lens to turn gateway traffic into agent improvements — ycombinator · 2026-10-01
- Meme: "Reviewing AI output" before pushing straight to prod — tlakomy · 2026-10-01
- A 5-second talking-head loop in 6GB VRAM: local graph, hosted lip-sync API — Extreme-Shock1930 · 2026-10-01
- My agents burned video credits on bad frames, so now they must ask before spending — coolxeo · 2026-10-01
- ComfyUI plugin Continuity update: in-shot screen replacement, image-to-3D, Qwen Image 2.1 — Fine_Rhubarb3786 · 2026-10-01