Meta-reasoning harness lifts GPT-5.5 to 71.5% on ProgramBench, beating Codex by 13.5 points
anirudhg9119 · x · 2026-10-02
An arXiv paper, Thinking Before Thinking: Scaling Agentic Inference Through Meta-Reasoning (authors include Ruslan Salakhutdinov, Jason Weston, Anirudh Goyal), introduces agentic meta-reasoning — an inference-time harness that makes execution control an explicit reasoning process for long-horizon agents.
- Architecture: Workers perform task-level computation while a controller consolidates progress, weighs options against the remaining budget, and dispatches work using compact persistent memory — carrying only a summary of the run between decisions instead of replaying full history.
- Results: On ProgramBench (long-horizon program reconstruction), meta-reasoning hits 71.5% with GPT-5.5 vs. 58.0% for Codex, and 67.2% with Opus 4.8 vs. 65.5% for Claude Code.
- Generalization: Gains of 3.6–4.2 points over direct-control baselines with identical compute across abstract reasoning, multi-domain long-horizon reasoning, and proof benchmarks.
The core claim: as agents tackle longer problems, controlling execution becomes a task in its own right and deserves a dedicated meta-reasoning layer.
More from coding & agent
- Reddit: vLLM and llama.cpp already ship training-free zero-shot classifiers via grammar-constrained output — Altruistic_Heat_9531 · 2026-10-02
- Dev Builds a Boxing Trainer App in One Shot with "5.5" — ZeroStateReflex · 2026-10-02
- Dev Wires Grok Bot and OpenAI Dot Into Voice Interfaces for a Hermes Agent Assistant — alexcovo_eth · 2026-10-02
- Translating an entire book with DeepSeek: pennies and under an hour, decent quality — teortaxesTex · 2026-10-02
- OmniSeek turns Omni-LLMs into agents that actively seek audio-visual evidence — Haibo Wang · 2026-10-02
- Microsoft's ActiveSaddler Uses Automated Curriculum Learning to Boost Agent Harnesses by 7.5 Points — microsoft · 2026-10-02