Arena now runs up to 600,000 E2B sandboxes a day for model evaluation
badphilosopher · x · 2026-10-02
E2B details its role powering Arena, which began in 2024 as a UC Berkeley Sky Computing Lab project for head-to-head model comparison and now evaluates models across agents, text, code, images, and video — running up to 600,000 E2B sandboxes daily.
Each sandbox is an isolated cloud computer where agents can write code, install dependencies, and work for hours on coding, research, reports, and presentations. Isolation keeps results trustworthy and secure: no session's code or files can reach another's and skew comparisons.
The company also notes that before GPT-5's public release, crowds rushed to Code Arena to try the model firsthand, with E2B's sandboxes serving as the backbone for that surge.
More from coding & agent
- Researcher shows a free open-source slide tool that beats PowerPoint and dodges AI design clichés — Afinetheorem · 2026-10-02
- Making a one-video history of the internet with Claude inside Cursor — prasenx · 2026-10-02
- Opus 5.5 turns Strudel live coding into an interactive jam session via MCP — repligate · 2026-10-02
- Skyvern 3.0 rebuilt from scratch hits SOTA 90.5% on Odyssey browser agent benchmark, stays open source — ycombinator · 2026-10-02
- How to test model switches for AI agents before shipping to catch silent tool-calling regressions — Fun_Employment6042 · 2026-10-02
- Databricks launches ai_decide, a sub-second structured decision AI function cheaper than LLMs — matei_zaharia · 2026-10-02