Agent Arena launches leaderboard grading models on millions of real agentic tasks
arena · x · 2026-09-11
Agent Arena launched a leaderboard that measures models on millions of real-world, long-horizon agentic tasks, giving them web search, filesystem and terminal tools to complete workflows like coding, slide decks and app building. It uses causal tracing to measure each model's net improvement over the average, with a Pareto frontier view.
More from coding & agent
- OpenAI appears to be quietly rolling out managed Agents on its platform — testingcatalog · 2026-09-11
- Microsoft Foundry swaps Model Router pool to GPT-5.6 variants and Claude Opus 4.8, ships Hosted Agents GA — WirelessLife · 2026-09-11
- Cua offers scaled computer-use agent fleets, but top frontier agent clears just 6 of 25 KiCad tasks — lucasmeijer · 2026-09-11
- Ouroboros traces program execution for LLM debugging, boosting small-model accuracy by up to 34 points — The_Homeless_God · 2026-09-11
- Coding agents have stopped writing code: they do the job themselves instead — doooyle · 2026-09-11
- Atto launches beta MCP server giving AI assistants a wallet with user-approved spending limits — Rotilho · 2026-09-11