LoopArena: even the best controller model hits just 24.69% managing coding agents

rohanpaul_ai · x · 2026-09-03

The LoopArena paper (arXiv 2608.28281) isolates the agent-management problem by fixing Qwen3.7-Plus as the coding Worker and swapping only the Controller. Across full 27-task runs, the best controller, GPT-5.5, reached just 24.69% Strict Success Rate, while simply restating the original goal each round scored 18.52% — identical to letting the Worker run uncontrolled. Useful control must react to evolving evidence, shifting the Worker between implementation, verification, recovery and stopping. Implication: benchmark the loop-managing model separately from the code-writing model.

Original post →

More from coding & agent

coding & agent channel →