Claude Fable 5.1 Tops Agent Arena With +15.8% Net Improvement, #1 in Praise Ratio
arena · x · 2026-09-06
The Agent Arena leaderboard now ranks Claude Fable 5.1 #1 overall, with a net improvement of +15.8%. Breakdown by signal:
- Praise vs. Complaint: #1 (+42.5%)
- Confirmed Success: #1 (+22.4%)
- Bash Recovery: #4 (+13.1%)
- Steerability: #28 (+0.6%) — its weakest area
- Tool Hallucination: no reported issues
The model clearly leads on user sentiment and task success, though steerability lags well behind.
Related event: Claude Fable 5.1 Tops Agent Arena Leaderboard(3 posts)→
More from Models
- Greenblatt walks back: OpenAI dropped reasoning=None as it's Pareto-dominated by low — RyanGreenblatt · 2026-09-06
- GPT-6 Astra reportedly almost never wrong on math, called most trustworthy model — gabrielchua · 2026-09-06
- User leaves GPT-6 Astra playing Unciv all night to test autonomous play — Angaisb_ · 2026-09-06
- Codex app users report Luna Max randomly disappearing while web version still works — Clear_Skye_ · 2026-09-06
- Where Fable dreams in opaque prose, Astra dreams in numbers — teortaxesTex · 2026-09-06
- Independent SpatialBench fully saturated by Astra, author declares LLM vision solved — pbaylies · 2026-09-06