Claude Fable 5.1 and Opus 5 Top AA-Briefcase Agentic Eval; GPT-6 Astra Gains ~85 Elo

ArtificialAnlys · x · 2026-09-05

Artificial Analysis published per-model results for AA-Briefcase, its private agentic knowledge-work evaluation in Index v4.2. Anthropic's Claude Fable 5.1 and Opus 5 lead, followed by GPT-6 Astra and Muse Spark 1.3, with GPT-6 Astra posting a 85 Elo jump over GPT-5.6 Sol. The eval runs models through multi-week projects with thousands of input files, combining rubric and pairwise grading on task success, analytical quality, and presentation quality.

Related event: Artificial Analysis Releases Intelligence Index v4.2; GPT-6 Astra Tops GDP.pdf Benchmark(6 posts)→

Original post →

More from Models

Models channel →