Agent Arena benchmarks models on millions of real-world long-horizon agentic tasks
arena · x · 2026-09-11
Agent Arena evaluates models on millions of real-world, long-horizon agentic tasks where models use web search, filesystem, and terminal tools to complete complex workflows.
Using causal tracing, its "net improvement" metric quantifies how much a model improves outcomes relative to the average model. The full leaderboard is available on its site.
More from coding & agent
- OpenAI ships GPT-Live prompting guide: copying your old prompts won't hit SOTA — craigsdennis · 2026-09-11
- Meta's Auto-RecSys runs autonomous research on industry-scale recommenders — omarsar0 · 2026-09-11
- Hamel Husain: The Revenge of the Data Scientist — AI Evals are the real job — HamelHusain · 2026-09-11
- One prompt builds a browser Rollercoaster Tycoon clone: "We are so EARLY" — _philschmid · 2026-09-11
- How to feed Claude Code local meeting notes without shipping transcripts to the cloud — Tight_Claim8869 · 2026-09-11
- Every Builds Personal Benchmarks for Each Employee — Finds a Smaller Model Beats the Big Ones — every · 2026-09-11