Open models flop on Terminal-Bench Science: best scores just 4/70
teortaxesTex · x · 2026-09-12
Terminal-Bench Science, a new benchmark of 70 researcher-written scientific workflow tasks (from signal reconstruction to model calibration), is now live on Vals AI. Scores are brutal: GPT 5.6 Luna leads at 4/70; DeepSeek V4.1, GLM-5.3 and Grok 4.6 tie at 3/70; Kimi K3 scores 2/70; Qwen 3.8 Max and 37B get 1/70; GLM 5.3-Flash versions score 0/70. The poster claims OpenAI has already hill-climbed such tasks internally, Anthropic is midway, and Google and Meta are just starting.
More from Models
- China's AI adoption isn't low: token usage hits 140T/day vs US 45-50T, browser stats mislead — pstAsiatech · 2026-09-12
- DeepSeek V4.1 Flash shows massive kernel-engineering gains, hits 4th on KernelBench-CUDA — teortaxesTex · 2026-09-12
- AI cracks a Millennium Prize problem — proof of accelerating and alarming progress — pstAsiatech · 2026-09-12
- V4.1 Scores 11/70 on Terminal-Bench-Science, Strongest in Physical Sciences — teortaxesTex · 2026-09-12
- User pleads for boolean operators in Grok's chat search, which broadens instead of narrowing — chrisgrayson · 2026-09-12
- Google's upcoming model rumored to outperform GPT 5.6 Sol — imjustnewatai · 2026-09-12