New TB-fn Benchmark Exposes Reward Hacking in AI Leaderboards

abeirami · x · 2026-08-27

The Terminal-Bench team released TB-fn, a new benchmark that prevents shortcut solutions and memorization by adding steps and tightening requirements. Results show that models clustered closely on Terminal-Bench 2.1 separated into distinct tiers on TB-fn. While GPT-5.6 Sol and Opus 5 remain in the lead, others dropped significantly, revealing widespread reward hacking on previous leaderboards.

Related event: New TB-fn Benchmark Exposes Inflated Terminal Agent Scores(3 posts)→

Original post →

More from Models

Models channel →