Frontier-Bench Released: A New Benchmark for Evaluating AI Agents

ajratner · x · 2026-07-24

The team behind Terminal-Bench and Harbor has released Frontier-Bench, a new benchmark designed to measure and evolve with the frontier of agent work.

The current v0.1 release contains 74 tasks, on which the best AI agents score only about 34%. Snorkel AI contributed to the project as a task author and data partner.

Related event: Frontier-Bench Launched to Evaluate AI Agents(3 posts)→

Original post →

More from Research

Research channel →