GPT-6 Astra and Claude Opus 5.5 Top New Terminal-Bench-Science Leaderboard

Artificial Analysis launched the Terminal-Bench-Science 0.1 leaderboard, built with Stanford researchers and the open-source community. GPT-6 Astra and Claude Opus 5.5 lead by roughly 20 points, with the benchmark clearly distinguishing models and reasoning levels—Opus 5.5 gained 38 points at max compute but at 5x the cost.

2026-09-25 ~ 2026-09-25 · 3 related posts