TerminalBench-Science selection: 70 tasks chosen from 920 proposals

BenBlaiszik · x · 2026-08-28

TerminalBench-Science set a high bar for task inclusion, selecting only 70 tasks for v0.1 from 920 proposals. Each task underwent domain review, technical review, and a final difficulty calibration. The author estimates over 80 hours of effort went into each task to ensure they truly challenge top-tier models.

Related event: Stanford-led Team Launches Terminal-Bench-Science, a Research-Agent Benchmark; Best Model Solves Only 30%(11 posts)→

Original post →

More from Research

Research channel →