GPT-6 Astra claims Terminal-Bench Science lead at 65.7%, 31.4 points clear of second place

DeryaTR_ · x · 2026-09-21

Vals AI has launched Terminal-Bench Science, a benchmark of 70 research workflow tasks written and reviewed by researchers, spanning signal reconstruction to model calibration, testing whether AI models can actually do scientific research.

According to the shared results, GPT-6 Astra leads at 65.7%, 31.4 points ahead of second-place Claude Fable 5.1 (34.3%). No other model clears 25%, and 12 of 24 models score 4.3% or lower. (Model names as relayed in the post, unverified.)

Original post →

More from Models

Models channel →