Stanford-led Team Launches Terminal-Bench-Science, a Research-Agent Benchmark; Best Model Solves Only 30%

The Stanford-led Terminal-Bench team, together with top universities worldwide, has released Terminal-Bench-Science (TBS) v0.1, a new benchmark for evaluating AI agents' ability to handle interdisciplinary scientific research workflows. The strongest current model, Claude Opus 5, passes only 30.0% of tasks, revealing a huge gap between frontier models and true research automation and providing a new yardstick for tracking progress in automating scientific workflows.

Confirmed

Why it matters

2026-08-28 ~ 2026-08-28 · 11 related posts

Primary sources