TerminalBench-Science sets high bar, slashing model scores by 10+ points

BenBlaiszik · x · 2026-08-28

Comparisons show models score significantly higher on Terminal-Bench 3.0 (e.g., Fable 5: 83.8%), while TBS pushes every model down by over 10 points. This difficulty was calibrated against the newest frontier models. The bar for inclusion was high: only 70 tasks made the cut from 920 proposals, with each requiring substantial effort.

Related event: Stanford-led Team Launches Terminal-Bench-Science, a Research-Agent Benchmark; Best Model Solves Only 30%(11 posts)→

Original post →

More from Research

Research channel →