Terminal-Bench-Science leaderboard launches with GPT-6 Astra at 63%

scaling01 · x · 2026-09-25

Artificial Analysis launched its leaderboard for Terminal-Bench-Science 0.1, a 70-task agentic benchmark built by Stanford researchers and the Terminal-Bench team covering life, physical, mathematical, engineering and earth sciences. GPT-6 Astra (max) leads at 63% and Claude Opus 5.5 (xhigh) at 62%, with agents graded pass/fail in sandboxed tasks via average pass@1 over 3 attempts.

Related event: GPT-6 Astra and Claude Opus 5.5 Top New Terminal-Bench-Science Leaderboard(3 posts)→

Original post →

More from Models

Models channel →