TB-fn benchmark reveals score drops for several models, exposing potential benchmark overfitting
burny_tech · x · 2026-08-28
Fidian released TB-fn, a variant of Terminal-Bench, revealing significant score drops for several models and exposing potential overfitting. Grok 4.5, Kimi K3, and GLM-5.2 dropped by 6-11 points, while only OpenAI and Anthropic models remained stable at the frontier. GLM-5.3 Flash leads open weights but its gap to Sol max widened from 1.0 to 8.6 points. The benchmark also showed a 41% average increase in per-attempt cost across all models.
Related event: New TB-fn Benchmark Exposes Suspected Benchmark Gaming in Terminal Models(5 posts)→
More from Models
- Users suspect DeepSeek V4 quality drop due to routing changes — l33thax0r_ · 2026-08-28
- PhoneLLM Alpha 1 Trends on Hugging Face: Voice Agent Model with MoE — pipecat-ai · 2026-08-28
- Why do research labs prefer Off-Policy Distillation for model improvement? — miifanboy · 2026-08-28
- 4x3090 Benchmarks: 27B Q8KXL at 32GB Beats Next IQ3 at 82GB on Both Speed and Security Tasks — Repulsive_Initial308 · 2026-08-28
- For Agent Memory, the Boring DeepSeek Non-Thinking Pass Wins on Speed and Accuracy — Brave_Pressure_9886 · 2026-08-28
- HALP: AI can know it's about to hallucinate before generating a single token — thisdudelikesAI · 2026-08-28