Stanford-led Team Launches Terminal-Bench-Science, a Research-Agent Benchmark; Best Model Solves Only 30%
The Stanford-led Terminal-Bench team, together with top universities worldwide, has released Terminal-Bench-Science (TBS) v0.1, a new benchmark for evaluating AI agents' ability to handle interdisciplinary scientific research workflows. The strongest current model, Claude Opus 5, passes only 30.0% of tasks, revealing a huge gap between frontier models and true research automation and providing a new yardstick for tracking progress in automating scientific workflows.
Confirmed
- v0.1 contains 70 tasks spanning life sciences, physics, mathematics, engineering, and earth sciences, built on real computational workflows and vetted through multi-stage expert review
- On overall pass rate, Claude Opus 5 (Claude Code) leads with 30.0% over GPT-5.6 Sol; by field, GPT-5.6 Sol wins in mathematics with 31.4%
- The 70 tasks were selected from 920 proposals, with each task going through three layers of review—domain reviewer, technical reviewer, and Bar Raiser; the authors estimate each task took over 80 hours to build
- Compared with Terminal-Bench 3.0, TBS is similarly effective at differentiating frontier models but lowers all model pass rates by more than 10 percentage points (e.g., Opus 5 drops from 43% to 30%; in contrast, the same models score significantly higher on Terminal-Bench 3.0, with Fable 5 reaching 83.8%), with difficulty calibrated specifically for the latest frontier models
- Contributing institutions include Stanford, MIT, Princeton, Tsinghua, Peking University, Oxford, and Cambridge; the project also has a large team of reviewers, science advisors, and AI research advisors from major AI labs, all publicly acknowledged
- Positioning: filling the gap in measuring research capability by testing agents on multidisciplinary research workflows and tracking the prerequisite capabilities needed for "decades or a century of progress" (BenBlaiszik relayed related commentary)
Why it matters
- Accelerating scientific workflows is seen as the biggest lever for advancing humanity; a dedicated benchmark for this capability was previously lacking, and TBS gives the community a trackable yardstick
- Difficulty is calibrated to the latest frontier models with discrimination comparable to existing benchmarks, making it effective for measuring future models' research capability
- Tasks were contributed by researchers worldwide through an open review process, with rigorous selection ensuring benchmark quality and saturation resistance
2026-08-28 ~ 2026-08-28 · 11 related posts
Primary sources
- TerminalBench-Science v0.1: Opus 5 leads overall benchmarks — BenBlaiszik ·
- TerminalBench-Science selection: 70 tasks chosen from 920 proposals — BenBlaiszik ·
- Terminal-Bench-Science lowers model pass rates by over 10% — sanmikoyejo ·
- Stanford releases Terminal-Bench-Science: AI agent benchmark for research workflows — BenBlaiszik · 2026-08-28
- TerminalBench-Science released: Opus 5 achieves 30% pass rate — BenBlaiszik · 2026-08-28
- [source] TerminalBench-Science selection: 70 tasks chosen from 920 proposals — BenBlaiszik · 2026-08-28
- TerminalBench-Science sets high bar, slashing model scores by 10+ points — BenBlaiszik · 2026-08-28
- [source] TerminalBench-Science v0.1: Opus 5 leads overall benchmarks — BenBlaiszik · 2026-08-28
- TerminalBench-Science thanks extensive reviewer and advisor team — BenBlaiszik · 2026-08-28
- Terminal-Bench-Science Aims to Fill the Measurement Gap in Science — BenBlaiszik · 2026-08-28
- Stanford Releases Terminal-Bench-Science: Claude Opus 5 Solves Only ~30% — BenBlaiszik · 2026-08-28
- 70 tasks selected from 920 proposals for benchmark — sanmikoyejo · 2026-08-28
- [source] Terminal-Bench-Science lowers model pass rates by over 10% — sanmikoyejo · 2026-08-28
- Terminal-Bench-Science built with global top university scientists — sanmikoyejo · 2026-08-28