Arena now runs up to 600,000 E2B sandboxes a day for model evaluation

badphilosopher · x · 2026-10-02

E2B details its role powering Arena, which began in 2024 as a UC Berkeley Sky Computing Lab project for head-to-head model comparison and now evaluates models across agents, text, code, images, and video — running up to 600,000 E2B sandboxes daily.

Each sandbox is an isolated cloud computer where agents can write code, install dependencies, and work for hours on coding, research, reports, and presentations. Isolation keeps results trustworthy and secure: no session's code or files can reach another's and skew comparisons.

The company also notes that before GPT-5's public release, crowds rushed to Code Arena to try the model firsthand, with E2B's sandboxes serving as the backbone for that surge.

Original post →

More from coding & agent

coding & agent channel →