Terminal-Bench-Science launches: Claude Opus 5 solves only ~30% of research tasks

BenBlaiszik · x · 2026-08-28

A Stanford-led community effort released Terminal-Bench-Science, a benchmark evaluating AI agents on research workflows across scientific domains. v0.1 ships with 70 tasks, built by the Terminal-Bench team with domain experts at research institutions worldwide.

Headline result: Claude Opus 5 solves only 30%. In a computational-chemistry task on X-ray diffraction phase analysis — a daily chore for thousands of chemists — top frontier models on max effort grappled for hours but couldn't clear the bar of expert scientific judgement. The team hopes the benchmark drives progress on frontier science.

Related event: Stanford-led Terminal-Bench-Science Benchmark Released; Top Model Passes Only 30%(17 posts)→

Original post →

More from Research

Research channel →