GPT-6 Astra and Opus 5.5 Lead Terminal-Bench-Science by ~20 Points
ArtificialAnlys · x · 2026-09-25
Artificial Analysis testing shows GPT-6 Astra and Claude Opus 5.5 lead Terminal-Bench-Science by 20 points over Fable 5.1, with the best non-frontier-lab model, Qwen3.8 Max (0902), at just 12%. The latest GPT and Claude releases improved scores while cutting per-task cost versus their predecessors.
Related event: GPT-6 Astra and Claude Opus 5.5 Top New Terminal-Bench-Science Leaderboard(3 posts)→
More from Models
- Hesamation jokes he's back in a 'toxic relationship' to try Claude Opus 5.5 — addyosmani · 2026-09-25
- Ethan Mollick: GPT-6 Astra beats Nethack on just its 3rd try — emollick · 2026-09-25
- AI models ran a vending machine business for a year: GPT-6 Sol turned $500 into $14,428 — 141_1337 · 2026-09-25
- Kevin Roose hands an AI agent $100 and a Kalshi account to test frontier models — MickeySteamboat · 2026-09-25
- Ex-Meta engineer benchmarks Jev vs GPT-nano: same score, 5x faster, 20% pricier — danielmckinn0n · 2026-09-25
- Altimeter CEO: OpenAI's Navier–Stokes-solving model withheld amid safety and govt scrutiny — rohanpaul_ai · 2026-09-25