TerminalBench Deep Dive: GPT-5.6 Leads, Many Models Drop in Rank

abeirami · x · 2026-08-27

Top 20 models from TerminalBench 2.1 were re-evaluated on the stricter TB-fn benchmark. Results show a widening performance gap, with GPT-5.6 Sol and Opus 5 maintaining their lead, while others see significant drops. TB-fn adds steps and removes shortcuts, increasing task difficulty by 38% on average. Pareto frontier analysis also shows GPT-5.6 Luna and DeepSeek V4 Flash outperforming ox-alpha in cost-effectiveness.

Related event: TB-fn Benchmark Exposes Inflated Terminal Model Scores(4 posts)→

Original post →

More from Models

Models channel →