GPT-6 Astra and Opus 5.5 Lead Terminal-Bench-Science by ~20 Points

ArtificialAnlys · x · 2026-09-25

Artificial Analysis testing shows GPT-6 Astra and Claude Opus 5.5 lead Terminal-Bench-Science by 20 points over Fable 5.1, with the best non-frontier-lab model, Qwen3.8 Max (0902), at just 12%. The latest GPT and Claude releases improved scores while cutting per-task cost versus their predecessors.

Related event: GPT-6 Astra and Claude Opus 5.5 Top New Terminal-Bench-Science Leaderboard(3 posts)→

Original post →

More from Models

Models channel →