New TB-fn Benchmark Exposes Model Performance Gaps on Terminal Tasks
abeirami · x · 2026-08-26
Researchers released TB-fn, a stricter benchmark that reworks Terminal Bench 2.1 tasks by adding steps, tightening requirements, and requiring complex algorithms to eliminate shortcut solutions. When tested on TB-fn, models that clustered together on the original leaderboard spread across performance tiers, revealing a wider gap between open-weight families and frontier models.
More from Research
- Ran Boltz-2 100 million times to simulate cell biology — dom_beaini · 2026-08-26
- RMSNorm projects activations to a hypersphere, doesn't solve interpretability — _xjdr · 2026-08-26
- LangChain open-sources WikiBench to measure how much codebase wikis help coding agents — LangChain · 2026-08-26
- LpWM Research: Sparse Representations Make Latent Dynamics Easier to Model — randall_balestr · 2026-08-26
- Perplexity Reveals Dream Agents for Continuous Self-Improvement — perplexity_ai · 2026-08-26
- AWS Paper Reveals the 'Handoff Tax' in AI Agent Model Escalation — omarsar0 · 2026-08-26