Claude Fable 5.1 leads GDPval-AA at 1853 Elo, edging Opus 5 within overlapping CIs
ArtificialAnlys · x · 2026-09-02
Artificial Analysis released results on two benchmarks run with its open-source agent harness, Stirrup. On GDPval-AA v2, Claude Fable 5.1 (max) leads at 1853 Elo over Claude Opus 5 (max) at 1824, though confidence intervals overlap.
On AA-Briefcase they are effectively tied (1694 vs 1685), with diverging sub-scores: Fable 5.1 wins on analytical quality (2025 vs 1980) but trails on presentation (1495 vs 1572). Both benchmarks test whether models produce accurate, well-presented professional outputs.
More from Models
- MSL Releases Muse Voice: SOTA Real-Time Transcription Model — bowenc0221 · 2026-09-02
- Fable 5.1 strong at biology but frequent refusals limit utility — kenbwork · 2026-09-02
- Anthropic allows Fable 5.1 for vuln scanning, but prompt triggers downgrade — AccBalanced · 2026-09-02
- ChatGPT Remains Sycophantic While Logged In; Free and Paid Versions Show Significant Differences — lilyraynyc · 2026-09-02
- Fable 5.1 system prompt leaked by jailbreaker Pliny within an hour of release — Polymarket · 2026-09-02
- Reddit Predicts Open-Weight Models Won't Match Fable 5.1 Until Late 2025 — power97992 · 2026-09-02