New TB-fn Benchmark Reveals Significant Drops in Terminal Model Rankings

abeirami · x · 2026-08-27

Questions about the actual capabilities of top models on the Terminal Bench 2.1 leaderboard led to the creation of the TB-fn benchmark. TB-fn reworks the same 89 tasks by adding steps, tightening requirements, or necessitating more complex algorithms, while removing shortcut solutions. Results show models clustered on TB-2.1 spread across more tiers on TB-fn, with GPT-5.6 Sol and Opus 5 remaining in the lead, while others drop significantly. Models work 38% harder on average in TB-fn.

Related event: New TB-fn Benchmark Exposes Inflated Terminal Agent Scores(3 posts)→

Original post →

More from Models

Models channel →