Stanford Releases Terminal-Bench-Science: Claude Opus 5 Solves Only ~30%

BenBlaiszik · x · 2026-08-28

Stanford-led community project Terminal-Bench-Science (v0.1) is released to evaluate AI agents on research workflows across life, physical, mathematical, and earth sciences. It features 70 tasks built from real computational workflows with a high-quality bar via expert review. Initial results show Claude Opus 5 solving only about 30% of the tasks.

Related event: Stanford-led Team Launches Terminal-Bench-Science, a Research-Agent Benchmark; Best Model Solves Only 30%(11 posts)→

Original post →

More from coding & agent

coding & agent channel →