Claude Fable 5.1 leads GDPval-AA at 1853 Elo, edging Opus 5 within overlapping CIs

ArtificialAnlys · x · 2026-09-02

Artificial Analysis released results on two benchmarks run with its open-source agent harness, Stirrup. On GDPval-AA v2, Claude Fable 5.1 (max) leads at 1853 Elo over Claude Opus 5 (max) at 1824, though confidence intervals overlap.

On AA-Briefcase they are effectively tied (1694 vs 1685), with diverging sub-scores: Fable 5.1 wins on analytical quality (2025 vs 1980) but trails on presentation (1495 vs 1572). Both benchmarks test whether models produce accurate, well-presented professional outputs.

Original post →

More from Models

Models channel →