New TB-fn Benchmark Exposes Inflated Terminal Agent Scores
The Terminal-Bench team released TB-fn, a stricter benchmark that removes shortcuts and memorized solutions by adding steps and complex algorithms, causing many top models on Terminal-Bench 2.1 to drop sharply in ranking and revealing inflated scores.
2026-08-26 ~ 2026-08-27 · 3 related posts
- New TB-fn Benchmark Exposes Model Performance Gaps on Terminal Tasks — abeirami · 2026-08-26
- New TB-fn Benchmark Reveals Significant Drops in Terminal Model Rankings — abeirami · 2026-08-27
- New TB-fn Benchmark Exposes Reward Hacking in AI Leaderboards — abeirami · 2026-08-27