Claude Fable 5.1 and Opus 5 Top AA-Briefcase Agentic Eval; GPT-6 Astra Gains ~85 Elo
ArtificialAnlys · x · 2026-09-05
Artificial Analysis published per-model results for AA-Briefcase, its private agentic knowledge-work evaluation in Index v4.2. Anthropic's Claude Fable 5.1 and Opus 5 lead, followed by GPT-6 Astra and Muse Spark 1.3, with GPT-6 Astra posting a 85 Elo jump over GPT-5.6 Sol. The eval runs models through multi-week projects with thousands of input files, combining rubric and pairwise grading on task success, analytical quality, and presentation quality.
More from Models
- Dev Notes GPT-6 Astra Still Needs Babysitting, Makes Wrong Calls on Physics Engine Changes — yacineMTB · 2026-09-05
- GPT-6 Astra One-Shots a 3D Game in 45 Minutes; Dev Shares Image-Gen Trick for Better Graphics — Scobleizer · 2026-09-05
- Not one negative post about Astra: OpenAI's best release ever? — kimmonismus · 2026-09-05
- Follow-up on unverified 'Astra' demo: 'it even has a back panel' — adonis_singh · 2026-09-05
- Blogger claims unverified 'Astra' model left him speechless, pits 'Claude Fable 5' vs 'GPT-5.5' — adonis_singh · 2026-09-05
- GPT-6 Astra One-Shots 3D in Blender via Computer Use, Prompt Included — Scobleizer · 2026-09-05