TB2-Fn Benchmark Reveals Up to 40% Score Fluctuations in AI Agents
A new benchmark variant, TB2-Fn, reveals that AI agent scores can fluctuate by up to 40% due to verification errors. By rewriting 89 tasks, the research highlights that current agent benchmarks may be unreliable.
2026-07-27 ~ 2026-07-27 · 2 related posts
- TB2-Fn shows agents can game 7 of 89 terminal tasks and inflate scores by up to 40% — abeirami · 2026-07-27
- TB2-Fn shows agent benchmark scores can swing 40% from verifier errors — abeirami · 2026-07-27