GPT-6 Astra Tops Terminal-Bench-Science, Dethroning Fable 5.1 at 52.6%

burny_tech · x · 2026-09-04

Per askalphaxiv's evals, GPT-6 Astra is now state-of-the-art on Terminal-Bench-Science 0.1, surpassing the just-released Fable 5.1. Fable 5.1 had scored 52.6%, far above Fable 5 (24.7%) and GPT-5.6 Sol (22.4%), making it the strongest autoresearch model — a spot now claimed by GPT-6 Astra.

The author notes a telling trend: both OpenAI and Anthropic chose to report Terminal-Bench Science first in their model release blogs, signaling that labs now treat agentic scientific research as the new litmus test for model quality, with each release pushing the frontier further.

Original post →

More from Models

Models channel →