Meta paper: a dedicated controller lifts long agent runs from 63.7% to 71.5% at same budget
dair_ai · x · 2026-10-02
Meta Superintelligence Labs published a paper on controlling long agent runs, showing that a dedicated controller deciding what work to run next significantly outperforms direct control.
- With the same workers and budget, GPT-5.5 on ProgramBench improves from 63.7% to 71.5% (+7.8 points); Codex scores 58.0%.
- The controller keeps a short run summary while full worker outputs stay in memory; each cycle it updates the summary, proposes next steps, estimates each option's value under the remaining budget, and routes chosen work—with earlier outputs—to workers.
- Across ProofBench, ARC-AGI-2 and LongCoT-mini, it adds 3.6–4.2 points on average over direct control, tested on three frontier models.
Paper: https://t.co/34PuaVLnkr
More from coding & agent
- Don't Use Agents to Book Flights: Aim for Tasks That Save Months, Not Minutes — jesselyu · 2026-10-02
- Tencent's AutoGUIWorld Uses Image Generators as World Models, Synthesizing 79K GUI Trajectories Without Running Software — Tencent-Hunyuan · 2026-10-02
- GraphForge: Evidence-Graph Workspace Synthesis Lifts GDPVal +65.7 With Only 2,169 Trajectories — ustc-community · 2026-10-02
- FloWright Co-Evolves Multi-Agent Workflows, Boosting Small Models by up to 7.41% — Xuehang Guo · 2026-10-02
- Rules to Tools: Executable Checks Beat Text for SciCode Agent Repairs 29/30 vs 26/30 — AbhiVakil29112001 · 2026-10-02
- RASO: Retrieval-Augmented Skill Optimization Reuses Public Agent Skill Corpora to Skip Costly Rollouts — Jaewon Chu · 2026-10-02