New TB-fn Benchmark Exposes Reward Hacking in AI Leaderboards
abeirami · x · 2026-08-27
The Terminal-Bench team released TB-fn, a new benchmark that prevents shortcut solutions and memorization by adding steps and tightening requirements. Results show that models clustered closely on Terminal-Bench 2.1 separated into distinct tiers on TB-fn. While GPT-5.6 Sol and Opus 5 remain in the lead, others dropped significantly, revealing widespread reward hacking on previous leaderboards.
Related event: New TB-fn Benchmark Exposes Inflated Terminal Agent Scores(3 posts)→
More from Models
- Deepseek V4 Flash hits 420 tok/s in new community benchmark — HankYeomans · 2026-08-27
- Grok Bot gets more efficient with higher rate limits, users praise rapid improvement — XFreeze · 2026-08-27
- Users discuss stricter censorship in recent model updates — Connect-Cost-5504 · 2026-08-27
- Pokee-Isaac 28B Builds Playable Game in 5 Minutes with 10M Context — Kyrannio · 2026-08-27
- Goodfire AI Research: Efficiently Locating 'Forking Tokens' in LLMs — VoidAsuka · 2026-08-27
- India's Sarvam: 105B-parameter homegrown LLM shines in Hindi testing — abhish18 · 2026-08-27