Frontier-Bench launches with 74 agent tasks and top systems scoring about 34%

ajratner · x · 2026-07-24

Frontier-Bench is live: a new ongoing benchmark for agentic work built by the team behind Terminal-Bench and Harbor.

The benchmark is designed to measure and evolve with the frontier of agent work rather than stay static. Version 0.1 includes 74 tasks, and the team says the best agents score about 34%.

The post also notes that SnorkelAI helped as a task author and data partner, contributing to benchmark-wide testing, corrections, and the category taxonomy.

Related event: Frontier-Bench v0.1 Released: Top Agents Score Only 34%(5 posts)→

Original post →

More from Research

Research channel →