Frontier models fail at counting letters: Astra hits 93% accuracy, Fable only 53%

maksym_andr · x · 2026-09-18

maksymandr ran letter-counting evaluations on frontier LLMs—a trivial task but out-of-distribution for models benchmaxxed on SWE-Bench/TerminalBench:

Follow-ups: harmless counting of random word sequences triggers Fable's bio safeguards, and Fable burns 3x+ more tokens than Astra on long passages while scoring worse.

Related event: Letter-counting tests expose Fable 5.1 false safety triggers and poor token efficiency(5 posts)→

Original post →

More from Models

Models channel →