Open models flop on Terminal-Bench Science: best scores just 4/70

teortaxesTex · x · 2026-09-12

Terminal-Bench Science, a new benchmark of 70 researcher-written scientific workflow tasks (from signal reconstruction to model calibration), is now live on Vals AI. Scores are brutal: GPT 5.6 Luna leads at 4/70; DeepSeek V4.1, GLM-5.3 and Grok 4.6 tie at 3/70; Kimi K3 scores 2/70; Qwen 3.8 Max and 37B get 1/70; GLM 5.3-Flash versions score 0/70. The poster claims OpenAI has already hill-climbed such tasks internally, Anthropic is midway, and Google and Meta are just starting.

Original post →

More from Models

Models channel →