PINNACLE: Claude Fable 5.1 halves agent failure rate but costs $2.46 per correct answer
ryanshrout · x · 2026-09-03
Two days after Signal65 launched its PINNACLE leaderboard, Anthropic shipped Claude Fable 5.1, and the numbers show why it exists:
- Failure rate: agent workflow failures drop from 14/100 (where Claude Opus 5 and GPT-5.6 Sol sit) to 7/100 — roughly half
- Hallucination: it invents an answer on 0.7% of unanswerable questions vs 7.6% for Opus 5 (1 in 140 vs 1 in 13)
- Why it matters: that gap moves agents from "a person checks everything" to "a person checks the exceptions"
- The catch: $2.46 per correct answer, 1.8x Opus 5, mostly from context the agent re-reads every step; the bill scales with volume with no knob to turn it down
Fewer failures or a smaller bill, per role, is the trade-off PINNACLE is built to surface.
More from Infra
- Broadcom CEO: Anthropic and OpenAI to become its No.1 and No.2 custom AI chip clients — dinabass · 2026-09-03
- Hock Tan says power availability dictates the exact timing of compute capacity deployment — BenBajarin · 2026-09-03
- Perplexity open-sources its Mac inference server optimized for Qwen 3.6 on Apple Silicon — Specter_Origin · 2026-09-03
- GLM-5.3-Flash hits 1,005 tok/s locally on dual RTX PRO 6000 Blackwell cards — BanghuaZ · 2026-09-03
- Nvidia now ~8% of S&P 500 market cap, worth 16.3% of US GDP at $5.3T — ivan_bezdomny · 2026-09-03
- Broadcom Q3 AI chip revenue hits $16.7B, up 221% YoY; guides $21.7B for Q4 — Beth_Kindig · 2026-09-03