LoopArena: even the best controller model hits just 24.69% managing coding agents
rohanpaul_ai · x · 2026-09-03
The LoopArena paper (arXiv 2608.28281) isolates the agent-management problem by fixing Qwen3.7-Plus as the coding Worker and swapping only the Controller. Across full 27-task runs, the best controller, GPT-5.5, reached just 24.69% Strict Success Rate, while simply restating the original goal each round scored 18.52% — identical to letting the Worker run uncontrolled. Useful control must react to evolving evidence, shifting the Worker between implementation, verification, recovery and stopping. Implication: benchmark the loop-managing model separately from the code-writing model.
More from coding & agent
- Firecrawl open-sources anydoc: Rust library converts Word, PPT, Excel, PDF to Markdown in milliseconds — tom_doerr · 2026-09-03
- Tsinghua's Frontis-MA1 pushes 35B model to 71.21% on MLE-Bench Lite for recursive self-improvement — ceciletamura · 2026-09-03
- Codex Remote syncs queued messages, no longer needs iOS app in foreground — Dimillian · 2026-09-03
- Fully automated tracker discovers new Hugging Face models within a day and estimates max tok/s — helloiamleonie · 2026-09-03
- Clearing ~900 accumulated chats made Codex feel instantly lighter — CtrlAltDwayne · 2026-09-03
- Radix sort in pure fragment shaders: 5B key/value pairs/sec with zero compute shaders — Michael_Moroz_ · 2026-09-03