New TB-fn Benchmark Exposes Model Performance Gaps on Terminal Tasks

abeirami · x · 2026-08-26

Researchers released TB-fn, a stricter benchmark that reworks Terminal Bench 2.1 tasks by adding steps, tightening requirements, and requiring complex algorithms to eliminate shortcut solutions. When tested on TB-fn, models that clustered together on the original leaderboard spread across performance tiers, revealing a wider gap between open-weight families and frontier models.

Original post →

More from Research

Research channel →