Stanford launches Terminal-Bench-Science: Claude Opus 5 solves only ~30%

ajratner · x · 2026-08-28

Stanford-led community effort releases Terminal-Bench-Science, a benchmark evaluating AI agents on real research workflows across physical, life, earth, and mathematical sciences. Built by the Terminal-Bench team with domain experts at research institutions worldwide.

v0.1 contains 70 tasks; Claude Opus 5 solves only 30%, making it hard for frontier models. Tasks are derived from real computational workflows and pass through multi-stage expert review. Snorkel AI partners via its Open Benchmarks Grants program, with the Laude Institute and Harbor framework teams contributing.

Related event: Stanford-led Terminal-Bench-Science Benchmark Released; Top Model Passes Only 30%(17 posts)→

Original post →

More from coding & agent

coding & agent channel →