Terminal-Bench-Science lowers model pass rates by over 10%

sanmikoyejo · x · 2026-08-28

Terminal-Bench-Science discriminates between frontier models as effectively as Terminal-Bench 3.0 but pushes pass rates down by more than 10 points for every model. Specific drops: Opus 5 (43%→30%), GPT-5.6 Sol (34%→22%), and Fable 5 (34%→21%). This indicates the benchmark's tasks present a higher challenge for AI agents.

Related event: Stanford-led Terminal-Bench-Science Launches; Claude Opus 5 Solves Just 30%(16 posts)→

Original post →

More from Research

Research channel →