GPT-6 Astra Tops Terminal-Bench-Science, Dethroning Fable 5.1 at 52.6%
burny_tech · x · 2026-09-04
Per askalphaxiv's evals, GPT-6 Astra is now state-of-the-art on Terminal-Bench-Science 0.1, surpassing the just-released Fable 5.1. Fable 5.1 had scored 52.6%, far above Fable 5 (24.7%) and GPT-5.6 Sol (22.4%), making it the strongest autoresearch model — a spot now claimed by GPT-6 Astra.
The author notes a telling trend: both OpenAI and Anthropic chose to report Terminal-Bench Science first in their model release blogs, signaling that labs now treat agentic scientific research as the new litmus test for model quality, with each release pushing the frontier further.
More from Models
- Matt Shumer reviews GPT-6 Astra: first model he trusts to run his inbox and business — mattshumer_ · 2026-09-04
- Researcher disputes OpenAI's claim Astra is its most aligned model: metrics may just hide reward hacking — connoraxiotes · 2026-09-04
- Qwen 3.8 27B vs 3.6: quality up 8% but runtime 5x longer and 4x more tokens — DerTomsn · 2026-09-04
- Leak claims GPT-6 Astra trained on 100,000+ GPUs at OpenAI's Stargate site — BLUECOW009 · 2026-09-04
- Gary Marcus on GPT-6 Astra: symbolic world models vindicated, but not AGI — Gary Marcus · 2026-09-04
- Users dispute credit burn; provider says KV cache was always on, scaling across providers — arthurcolle · 2026-09-04