Claude Fable 5 Breaks WeirdML Record
scaling01 · x · 2026-07-12
Claude Fable 5 (max) scored 91.9% on WeirdML, setting a new record. Key points: - If you combine the historical best scores per task for a theoretical SOTA, the total would be 93.5%, indicating Fable 5 in a single run generally approaches task-specific bests. - It achieved SOTA on 7 out of 17 tasks. - Its worst run was only 7% below the corresponding SOTA. - This run used only 2 trials per task, fewer than the common 5, making the result more impressive but also more likely to dodge a few bad samples.
More from Models
- Claude 20x users report sharply tighter limits and faster quota burn — MarcJSchmidt · 2026-07-21
- Cola launches July, the latest model jokingly billed as “second only to Fable” — oran_ge · 2026-07-21
- Kimi K3 looks stronger and about 5× cheaper on a frontend dashboard task — OwariDa · 2026-07-21
- Last Week in AI recap: Anthropic’s $65B round, IPO filing, and Microsoft’s MAI push — Last Week in AI · 2026-07-21
- A user says Claude 4.6 felt worse yesterday and asks whether model quality can drift over time — Rahios · 2026-07-21
- Kimi K3 hits 89.4% peak on software tasks while Fable 5 is slightly steadier — FinanceYF5 · 2026-07-21