LoopArena: Open Benchmark Tests Which LLMs Make the Best Runtime Controllers for Coding Agents

PepsiBetter · reddit · 2026-09-07

LoopArena is an open benchmark evaluating models as runtime Controllers for long-running coding-agent systems—deciding next steps, verification, and stopping—while holding the Worker, tools, and budgets fixed. Three settings range from execution-validated next-step decisions (Type I) to full-task control from original states (Type III). Initial five-model panel shows the best Type III strict success rate is just 24.69%; Type II cuts estimated inference cost by 64.4% with similar Controller rankings. Data, code, and protocol are open-sourced.

Original post →

More from coding & agent

coding & agent channel →