Sonnet 5.5 jumps 50 points to 64% on Terminal-Bench 4.0, topping Opus 5.5

ArtificialAnlys · x · 2026-09-29

Artificial Analysis shared more Sonnet 5.5 results: 64% on Terminal-Bench 4.0, a 50-point jump over Sonnet 5 (max) and slightly above Opus 5.5 and GPT-6 Astra (xhigh) at 60%. On the new Terminal-Bench-Science leaderboard—agentic terminal use for realistic scientific research workflows—it scores 53%, behind only GPT-6 Astra and Opus 5.5. That benchmark isn't yet in the Intelligence Index.

Related event: Sonnet 5.5 Nearly Matches Opus 5.5 but Sets Token Consumption Record(12 posts)→

Original post →

More from Models

Models channel →