TerminalBench Deep Dive: GPT-5.6 Leads, Many Models Drop in Rank
abeirami · x · 2026-08-27
Top 20 models from TerminalBench 2.1 were re-evaluated on the stricter TB-fn benchmark. Results show a widening performance gap, with GPT-5.6 Sol and Opus 5 maintaining their lead, while others see significant drops. TB-fn adds steps and removes shortcuts, increasing task difficulty by 38% on average. Pareto frontier analysis also shows GPT-5.6 Luna and DeepSeek V4 Flash outperforming ox-alpha in cost-effectiveness.
Related event: TB-fn Benchmark Exposes Inflated Terminal Model Scores(4 posts)→
More from Models
- Zai launches GLM-5.3 Flash on Cloudflare: First multimodal GLM with 1M context — michellechen · 2026-08-27
- Anthropic reportedly releasing Fable 5.1, claimed 3-4 months ahead — bindureddy · 2026-08-27
- Google's new Gemma models breeze through Google's own reCAPTCHA v2 — Hour-Wish8158 · 2026-08-27
- Zhipu GLM-5.3-Flash launches on OpenRouter with 1M-token context — AccBalanced · 2026-08-27
- GLM 5.3 Flash now matches Sol 5.6 (Max) on Artificial Analysis' Agentic Index — PilgrimofHaqq2 · 2026-08-27
- AI enters adolescence: Small models beating large ones in specific domains — DhruvBatra_ · 2026-08-27