Terminal-Bench Team Launches Frontier-Bench for Agent Evaluation

BenBlaiszik · x · 2026-07-24

The team behind Terminal-Bench released Frontier-Bench, a new agentic benchmark designed to continuously evolve alongside rapidly advancing frontier models, preventing tests from becoming obsolete.

The current v0.1 release features 74 tasks, with top-scoring agents achieving around a 34% success rate. Aside from software engineering (SWE), science tasks make up the second-largest category, aiming to give science as much attention as coding in agentic evaluations.

Related event: Frontier-Bench v0.1 Released: Top Agents Score Only 34%(5 posts)→

Original post →

More from coding & agent

coding & agent channel →