Terminal-Bench Team Launches Frontier-Bench for Agent Evaluation
BenBlaiszik · x · 2026-07-24
The team behind Terminal-Bench released Frontier-Bench, a new agentic benchmark designed to continuously evolve alongside rapidly advancing frontier models, preventing tests from becoming obsolete.
The current v0.1 release features 74 tasks, with top-scoring agents achieving around a 34% success rate. Aside from software engineering (SWE), science tasks make up the second-largest category, aiming to give science as much attention as coding in agentic evaluations.
Related event: Frontier-Bench v0.1 Released: Top Agents Score Only 34%(5 posts)→
More from coding & agent
- The browser main thread is expensive: a practical guide to JavaScript and CSS animation cost — jh3yy · 2026-09-11
- Inspired by OpenAI's 10,000-agent run, dev open-sources a crowdsourced agent problem-solving platform — Benjaminsen · 2026-09-11
- Lucid: open-source Mac app keeps your laptop awake only while AI agents run — Pitiful_Hedgehog_600 · 2026-09-11
- banteg's snail project crowdsources AI agents to finish matching Snail Mail's 20 remaining functions — banteg · 2026-09-11
- Alex Townsend posts 200 open problems in numerical linear algebra for humans and AI agents — IgorCarron · 2026-09-11
- Kimi K2.8 Preview rolls out: near-K3 coding performance, 1M context for all tiers — teortaxesTex · 2026-09-11