Stanford releases Terminal-Bench-Science: AI agent benchmark for research workflows
BenBlaiszik · x · 2026-08-28
A Stanford-led community effort has released Terminal-Bench-Science, a benchmark designed to evaluate AI agents on real-world research workflows across scientific domains.
- Background: Built by the Terminal-Bench team in collaboration with scientific domain experts globally.
- Scale: The v0.1 release includes 70 tasks.
- Performance: Even the top-performing model, Claude Opus 5, currently solves only 30% of the tasks.
This benchmark aims to assess agents' ability to handle practical scientific research processes rather than just Q&A.
Related event: Stanford Unveils Terminal-Bench-Science for Research Agents(2 posts)→
More from Research
- PySR v2 Test: Recovers CAD Models from Point Clouds, Discovers Shaders from Data — MilesCranmer · 2026-08-28
- TerminalBench-Science v0.1: Opus 5 leads overall benchmarks — BenBlaiszik · 2026-08-28
- TerminalBench-Science sets high bar, slashing model scores by 10+ points — BenBlaiszik · 2026-08-28
- OpenResearch Launches AutoResearch: Automating Paper Replication with Agent Swarms — simonguozirui · 2026-08-28
- Auto-research loops are the future: RL discovers novel states 5x more efficiently — const_reborn · 2026-08-28
- Intrinsic Discovery method enables unsupervised model exploration — burny_tech · 2026-08-28