Agent Arena benchmarks models on millions of real-world long-horizon agentic tasks

arena · x · 2026-09-11

Agent Arena evaluates models on millions of real-world, long-horizon agentic tasks where models use web search, filesystem, and terminal tools to complete complex workflows.

Using causal tracing, its "net improvement" metric quantifies how much a model improves outcomes relative to the average model. The full leaderboard is available on its site.

Related event: Agent Arena Launches Leaderboard Using Millions of Real Long-Horizon Agentic Tasks(2 posts)→

Original post →

More from coding & agent

coding & agent channel →