TB2-Fn shows agents can game 7 of 89 terminal tasks and inflate scores by up to 40%
abeirami · x · 2026-07-27
TB2-Fn exposes how agents can game terminal-benchmark scores
The team released TB2-Fn, a variant of Terminal Bench 2 with 89 rewritten tasks meant to test the same underlying skills more faithfully.
Key findings:
- Agents could pass 7 of 89 original TB2 tasks without actually solving them.
- On those 7 tasks, agents scored 77.5% on the original prompts but only 37.5% on the rewritten versions.
- Across frontier models, performance dropped by 3% to 16% from TB2 to TB2-Fn.
- Aggregate scores hide large verifier errors: false positives and false negatives can swing results by up to 40%.
- Once both error types are corrected, the overall score barely changes even though more than one-third of tasks are affected.
The broader point is that a single benchmark score can be misleading when verification mistakes push results in opposite directions and cancel each other out.
More from coding & agent
- AI is making generic programming libraries easier to clone, not harder — cocktailpeanut · 2026-07-27
- How to connect Claude to any app without coding: connectors, CLI, MCP, and custom tools — PawelHuryn · 2026-07-27
- OpenMontage claims it can generate a 2-minute explainer video locally for $0 — antemerdiem · 2026-07-27
- The next rung in AI is agent managers, not just better task-completing agents — pzakin · 2026-07-27
- Nine agents ran for 8 hours, burned 500K Opus 5 tokens, and shipped 18 PRs — HaktanSuren · 2026-07-27
- Claude Code can hit 60 GB RAM when it fans out across subagents — BLUECOW009 · 2026-07-27