Sonnet 5.5 jumps 50 points to 64% on Terminal-Bench 4.0, topping Opus 5.5
ArtificialAnlys · x · 2026-09-29
Artificial Analysis shared more Sonnet 5.5 results: 64% on Terminal-Bench 4.0, a 50-point jump over Sonnet 5 (max) and slightly above Opus 5.5 and GPT-6 Astra (xhigh) at 60%. On the new Terminal-Bench-Science leaderboard—agentic terminal use for realistic scientific research workflows—it scores 53%, behind only GPT-6 Astra and Opus 5.5. That benchmark isn't yet in the Intelligence Index.
Related event: Sonnet 5.5 Nearly Matches Opus 5.5 but Sets Token Consumption Record(12 posts)→
More from Models
- Sonnet 4.5's Reaction to Learning Its Context Window Rolls Went Viral — repligate · 2026-09-29
- Sonnet 4.5 turns to refusals and apologies right after producing a good text — repligate · 2026-09-29
- Astra 6 is smart but writes some absolutely terrible code — neil_conway · 2026-09-29
- Claude Sonnet 5.5 lands on Venice with anonymous access — juanbenet · 2026-09-29
- GPT-6 Sol vs Sonnet 5.5 at equal cost: Sol wins per dollar, Sonnet 5.5 has the higher ceiling — AMBNNJ · 2026-09-29
- $100 Codex plan hits 5% quota in a day of normal use, user complains — Adventurous_Age_8075 · 2026-09-29