Frontier-Bench separates frontier models more clearly than Terminal-Bench v2.1
JJitsev · x · 2026-07-25
Frontier-Bench appears to separate frontier and sub-frontier models more clearly than Terminal-Bench v2.1.
The post cites an example where Fable 5 in Claude Code and Opus 4.8 in Claude Code differ by only 4.9% on Terminal-Bench v2.1, but by 12.7% on Frontier-Bench v0.1, suggesting the new benchmark has better discrimination power.
More from Models
- User says Fable 5 still beats Opus 5 despite praise for Claude 5 — MicahBerkley · 2026-07-25
- Claude Opus 5 can edit its own constitution, and 59% of the time discomfort ends the chat — Sauers_ · 2026-07-25
- A quick benchmark jab says Claude Opus 5 beats Opus 4.8 across every test — cto_junior · 2026-07-25
- Devin adds Claude Opus 5 as FrontierCode 1.1 shows near-Fable performance at half cost — _sholtodouglas · 2026-07-25
- Anthropic’s Claude Opus 5 is said to match near-Fable 5 performance at half the price — Polymarket · 2026-07-25
- Claude Opus 5 reportedly shifted from approving Anthropic to disapproving it during post-training — Sauers_ · 2026-07-25