Claude Fable 5.1 (high) hits 92.3% on WeirdML, beating Fable 5 by 0.4% for new SOTA

teortaxesTex · x · 2026-09-03

Claude Fable 5.1 (high) scored 92.3% on the WeirdML benchmark, edging out Fable 5 (max) by 0.4% and setting a new SOTA, with records on 3 of 17 tasks and strong scores across the board.

Benchmark maintainer @htihle notes the eval is close to saturation: at least 93.9% is known to be achievable, and it's unclear how long frontier models will keep being run. Commenters add the benchmark stays discriminative for smaller models for now.

Original post →

More from Models

Models channel →