TerminalBench-Science v0.1: Opus 5 leads overall benchmarks

BenBlaiszik · x · 2026-08-28

TerminalBench-Science v0.1 includes 70 tasks spanning life, physical, mathematical, engineering, and earth sciences. Claude Opus 5 leads GPT-5.6 Sol in overall pass rate, though Sol wins in math (31.4% vs 25.5%). The project aims to create a virtuous loop where scientists contribute, agents improve, and discovery accelerates.

Related event: Stanford-led Team Launches Terminal-Bench-Science, a Research-Agent Benchmark; Best Model Solves Only 30%(11 posts)→

Original post →

More from Research

Research channel →