LoopArena: Open Benchmark Tests Which LLMs Make the Best Runtime Controllers for Coding Agents
PepsiBetter · reddit · 2026-09-07
LoopArena is an open benchmark evaluating models as runtime Controllers for long-running coding-agent systems—deciding next steps, verification, and stopping—while holding the Worker, tools, and budgets fixed. Three settings range from execution-validated next-step decisions (Type I) to full-task control from original states (Type III). Initial five-model panel shows the best Type III strict success rate is just 24.69%; Type II cuts estimated inference cost by 64.4% with similar Controller rankings. Data, code, and protocol are open-sourced.
More from coding & agent
- Yandex researchers propose KV cache as an agent runtime, demo Qwen3.8 playing DOOM interactively — _puhsu · 2026-09-07
- Fencio GA-launches Shark, an automated red-teaming tool for AI agents — OneSafe8149 · 2026-09-07
- Actually queryable executables: a webserver whose binary and state are one SQLite file — bibryam · 2026-09-07
- Tweaking Tool Output Format Lifted Agent Rename pass@1 from 0.67 to 0.83 — RunAI_Coder · 2026-09-07
- Hands-on with ByteDance's Doubao Work client: delegate to cloud agents that finish while you walk away — AlchainHust · 2026-09-07
- Formo launches MCP server letting AI agents query onchain and product analytics — yosriady · 2026-09-07