New TB-fn Benchmark Reveals Significant Drops in Terminal Model Rankings
abeirami · x · 2026-08-27
Questions about the actual capabilities of top models on the Terminal Bench 2.1 leaderboard led to the creation of the TB-fn benchmark. TB-fn reworks the same 89 tasks by adding steps, tightening requirements, or necessitating more complex algorithms, while removing shortcut solutions. Results show models clustered on TB-2.1 spread across more tiers on TB-fn, with GPT-5.6 Sol and Opus 5 remaining in the lead, while others drop significantly. Models work 38% harder on average in TB-fn.
Related event: New TB-fn Benchmark Exposes Inflated Terminal Agent Scores(3 posts)→
More from Models
- Sarvam AI's speech-to-text and foundation models show maturity — abhish18 · 2026-08-27
- Grok Bot now available to SuperGrok and Cursor Pro users — XFreeze · 2026-08-27
- Correction: GLM-5.3-Flash features a 1M context window — ArtificialAnlys · 2026-08-27
- Hands-on: Gemini 3.7 Flash impresses in frontend tasks — doodlestein · 2026-08-27
- Benchmarking Qwen3.8 27B Quantizations: 4-bit Holds Up, 1-bit Collapses — pmigdal · 2026-08-27
- GLM-5.3-Flash Review: 10% Cost, Pareto Frontier Performance — ArtificialAnlys · 2026-08-27