Terminal-Bench-Science leaderboard launches with GPT-6 Astra at 63%
scaling01 · x · 2026-09-25
Artificial Analysis launched its leaderboard for Terminal-Bench-Science 0.1, a 70-task agentic benchmark built by Stanford researchers and the Terminal-Bench team covering life, physical, mathematical, engineering and earth sciences. GPT-6 Astra (max) leads at 63% and Claude Opus 5.5 (xhigh) at 62%, with agents graded pass/fail in sandboxed tasks via average pass@1 over 3 attempts.
Related event: GPT-6 Astra and Claude Opus 5.5 Top New Terminal-Bench-Science Leaderboard(3 posts)→
More from Models
- Is there any LLM whose training data is fully auditable and un-stolen? — LuCiAnO241 · 2026-09-25
- Can a small local LLM with internet access rival a larger model? — mototuneup · 2026-09-25
- Math Benchmark: Astra Dominates, Nothing Below Fable 5.1 Is Competitive — teortaxesTex · 2026-09-25
- Someone ran tests across all Claude models and published the results — repligate · 2026-09-25
- Open-Sourced 3D Pelican Bike Game: One Prompt, Zero Hand Edits, Prompt Included — EricBuess · 2026-09-25
- One Lazy Prompt to Claude Opus 5.5 Yields a Full 3D Storybook Farm Game — EricBuess · 2026-09-25