GPT-6 Astra and Claude Opus 5.5 Top New Terminal-Bench-Science Leaderboard
Artificial Analysis launched the Terminal-Bench-Science 0.1 leaderboard, built with Stanford researchers and the open-source community. GPT-6 Astra and Claude Opus 5.5 lead by roughly 20 points, with the benchmark clearly distinguishing models and reasoning levels—Opus 5.5 gained 38 points at max compute but at 5x the cost.
2026-09-25 ~ 2026-09-25 · 3 related posts
- Claude Opus 5.5 Gains 38 Points on Terminal-Bench-Science at 5x the Cost — ArtificialAnlys · 2026-09-25
- GPT-6 Astra and Opus 5.5 Lead Terminal-Bench-Science by ~20 Points — ArtificialAnlys · 2026-09-25
- Terminal-Bench-Science leaderboard launches with GPT-6 Astra at 63% — scaling01 · 2026-09-25