TB2-Fn shows agent benchmark scores can swing 40% from verifier errors
abeirami · x · 2026-07-27
TB2-Fn is a rewritten variant of Terminal Bench 2 with 89 tasks redesigned to test the same underlying agent skills more reliably.
- The authors found that agents could pass 7 of the original TB2 tasks without actually solving them.
- Across those 7 tasks, agents scored 77.5% on the originals versus 37.5% on the rewritten versions.
- Overall scores dropped by 3%–16% across frontier models when moving from TB2 to TB2-Fn.
- The bigger takeaway is that verifier false positives and false negatives can swing scores by up to 40%, and those errors can cancel each other out in aggregate.
- The paper argues that aggregate benchmark numbers can hide substantial task-level distortions.
Related event: TB2-Fn Benchmark Reveals Up to 40% Score Fluctuations in AI Agents(2 posts)→
More from coding & agent
- Agent loops get pricey because every call repays the whole conversation history — Future_AGI · 2026-07-27
- Claude Code experiment says Pangram's AI detector can be beaten after 17 rewrites — _akpiper · 2026-07-27
- NEO for VS Code and Cursor adds @ file tagging to streamline coding context — Nilofer_tweets · 2026-07-27
- A Stanford AI engineer launches a free 10-day AI basics series — HarperSCarroll · 2026-07-27
- Claude workflow can be deployed to run automatically every morning at 7 — TawohAwa · 2026-07-27
- Keel-opencore runs a bounded agent swarm for about 15 days of self-improvement — Efficient-Cap6662 · 2026-07-27