Terminal-Bench 3.0: Top Model Score Plummets to 43.5% as Agents Fail to Fix Root Causes

ajratner · x · 2026-08-25

Snorkel AI released a deep dive into Terminal-Bench 3.0. The new benchmark significantly raised the difficulty, causing the top model's score to plummet from 84% in version 2.1 to just 43.5% in version 3.0. The analysis highlights two critical failure patterns in current AI agents:

These cases reveal a significant gap in agent reasoning for complex, multi-component real-world engineering tasks.

Original post →

More from coding & agent

coding & agent channel →