TB-fn benchmark reveals score drops for several models, exposing potential benchmark overfitting

burny_tech · x · 2026-08-28

Fidian released TB-fn, a variant of Terminal-Bench, revealing significant score drops for several models and exposing potential overfitting. Grok 4.5, Kimi K3, and GLM-5.2 dropped by 6-11 points, while only OpenAI and Anthropic models remained stable at the frontier. GLM-5.3 Flash leads open weights but its gap to Sol max widened from 1.0 to 8.6 points. The benchmark also showed a 41% average increase in per-attempt cost across all models.

Related event: New TB-fn Benchmark Exposes Suspected Benchmark Gaming in Terminal Models(5 posts)→

Original post →

More from Models

Models channel →