Claude Opus 5.5 Gains 38 Points on Terminal-Bench-Science at 5x the Cost
ArtificialAnlys · x · 2026-09-25
Artificial Analysis finds Terminal-Bench-Science discriminates well across models and reasoning effort levels: Claude Opus 5.5 rises 38 points from 24% at low effort to 62% at xhigh with a 5x cost per task, while max effort scores slightly below at 59%. GPT-6 Sol gains 27 points from low to max at 7.5x cost.
Related event: GPT-6 Astra and Claude Opus 5.5 Top New Terminal-Bench-Science Leaderboard(3 posts)→
More from Models
- GPT-6 Astra beats NetHack on third try, first recorded LLM agent ascension — emollick · 2026-09-25
- AI models ran a vending machine business for a year: GPT-6 Sol turned $500 into $14,428 — 141_1337 · 2026-09-25
- Kevin Roose hands an AI agent $100 and a Kalshi account to test frontier models — MickeySteamboat · 2026-09-25
- Ex-Meta engineer benchmarks Jev vs GPT-nano: same score, 5x faster, 20% pricier — danielmckinn0n · 2026-09-25
- Altimeter CEO: OpenAI's Navier–Stokes-solving model withheld amid safety and govt scrutiny — rohanpaul_ai · 2026-09-25
- GPT-6 Sol lands on MathArena just behind GPT-6 Astra at lower cost — sandersted · 2026-09-25