GPT-6 Astra claims Terminal-Bench Science lead at 65.7%, 31.4 points clear of second place
DeryaTR_ · x · 2026-09-21
Vals AI has launched Terminal-Bench Science, a benchmark of 70 research workflow tasks written and reviewed by researchers, spanning signal reconstruction to model calibration, testing whether AI models can actually do scientific research.
According to the shared results, GPT-6 Astra leads at 65.7%, 31.4 points ahead of second-place Claude Fable 5.1 (34.3%). No other model clears 25%, and 12 of 24 models score 4.3% or lower. (Model names as relayed in the post, unverified.)
More from Models
- JEV opens to all with $5 free credits; $0.042 input and free output pricing — op7418 · 2026-09-21
- Inception CEO Stefano Ermon bets on diffusion LLMs that generate tokens in parallel — saranormous · 2026-09-21
- Testing decision model Jev: add "none of these" options and never do algebra across questions — colinmcnamara · 2026-09-21
- Jev's calibration error measured at ~0.09: right just as often, still off about how sure — colinmcnamara · 2026-09-21
- Jev runs at 143ms per question, hits 94% on sentiment with zero-shot prompts — colinmcnamara · 2026-09-21
- TypeSafe's Jev: a model that only makes typed decisions and never writes prose — colinmcnamara · 2026-09-21