Claude Fable 5.1 (high) hits 92.3% on WeirdML, beating Fable 5 by 0.4% for new SOTA
teortaxesTex · x · 2026-09-03
Claude Fable 5.1 (high) scored 92.3% on the WeirdML benchmark, edging out Fable 5 (max) by 0.4% and setting a new SOTA, with records on 3 of 17 tasks and strong scores across the board.
Benchmark maintainer @htihle notes the eval is close to saturation: at least 93.9% is known to be achievable, and it's unclear how long frontier models will keep being run. Commenters add the benchmark stays discriminative for smaller models for now.
More from Models
- Tencent Hunyuan 770B compressed from ~1.5TB to ~214GiB with mixed quantization — Aiden_Tech_Ai · 2026-09-03
- Reddit user bails on GPT-5.6 over 'defensive slop' behavior and broken cron jobs — PersimmonLevel3500 · 2026-09-03
- 33.3% of ImageNet images contain more than one class — the answer key itself is flawed — RexDouglass · 2026-09-03
- Cohere Labs head: multilingual AI's 'longitude problem' — scale isn't enough — Cohere_Labs · 2026-09-03
- Gemini keeps generating images even when users explicitly say stop — a tool-calling UX bug report — JacketMajor6561 · 2026-09-03
- antirez: Zero Out Steering When Possible—It Always Adds Distortion — antirez · 2026-09-03