New TB2-Fn Benchmark Released: Fixes 38% Flawed Tasks in Original Terminal-Bench
abeirami · x · 2026-07-28
Fidian AI has released Terminal-Bench 2 Fidian Edition (TB2-Fn), a completely rebuilt variant of the Terminal-Bench agent benchmark.
- Fixing Flaws: An audit of the original TB2 revealed that roughly 38% of the 89 tasks had systematic verification issues. This included 7 score-inflating tasks (due to loopholes) and 27 score-deflating tasks (caused by ambiguous instructions or broken setups).
- Increased Difficulty: The new edition rewrites instructions, environments, and verification logic, pushing the boundaries for frontier AI models.
- Measuring True Skills: By varying the surface form of the tasks, TB2-Fn helps identify test-set fit. An agent passing both benchmarks demonstrates actual underlying capability, whereas a gap indicates overfitting.
Related event: New TB2-Fn Benchmark Reveals Up to 40% Error in Terminal Agent Evaluations(3 posts)→
More from coding & agent
- New study says LLM agent skills create a 59% regression tax across 5,832 runs — TheTuringPost · 2026-07-28
- Arena’s WebDev board puts Claude Opus 5 Max ahead of Kimi K3 Max — arena · 2026-07-28
- Gauntlet Loop splits goals across builder agents and a ruthless blind critic — mattshumer_ · 2026-07-28
- Baseten adds day-0 API support for Moonshot’s Kimi K3 on vLLM — AccBalanced · 2026-07-28
- Gauntlet Loops are becoming the author’s default workflow for almost every project — mattshumer_ · 2026-07-28
- MCP builders debate when real usage turns into stars, reviews and paid traction — ResponsibleOne6307 · 2026-07-28