Claude Opus 5.5 Gains 38 Points on Terminal-Bench-Science at 5x the Cost

ArtificialAnlys · x · 2026-09-25

Artificial Analysis finds Terminal-Bench-Science discriminates well across models and reasoning effort levels: Claude Opus 5.5 rises 38 points from 24% at low effort to 62% at xhigh with a 5x cost per task, while max effort scores slightly below at 59%. GPT-6 Sol gains 27 points from low to max at 7.5x cost.

Related event: GPT-6 Astra and Claude Opus 5.5 Top New Terminal-Bench-Science Leaderboard(3 posts)→

Original post →

More from Models

Models channel →