Frontier models fail at counting letters: Astra hits 93% accuracy, Fable only 53%
maksym_andr · x · 2026-09-18
maksymandr ran letter-counting evaluations on frontier LLMs—a trivial task but out-of-distribution for models benchmaxxed on SWE-Bench/TerminalBench:
- At low reasoning effort, GPT-6-Astra maintains 93% accuracy on 1280-character passages; Fable 5.1 fails much more often at 53%.
- Fable's accuracy curve is non-monotonic: below 320 characters it uses zero thinking tokens even though reasoning would help, exposing a failure of adaptive thinking.
- Motivation: probing no/low-CoT performance and whether models could execute complex plans (e.g. scheming) without verbalizing them in CoT.
Follow-ups: harmless counting of random word sequences triggers Fable's bio safeguards, and Fable burns 3x+ more tokens than Astra on long passages while scoring worse.
More from Models
- 105 planted bugs benchmark: Unbiased's Pareto scores 30.7 for just $4.81 — PawelHuryn · 2026-09-18
- Jason Wei's Stanford talk: intelligence is becoming a commodity as adaptive compute takes off — dotey · 2026-09-18
- GPT-6-Astra beats Fable-5.1 at RollerCoaster Tycoon 2 in 3 hours, using 5x fewer tokens — scaling01 · 2026-09-18
- RL agents invent their own diagnostic renderings to ground code understanding, sparking RL scaling optimism — teortaxesTex · 2026-09-18
- Dev discovers Codex security hardening switched persistent agent sessions to per-message instances — RileyRalmuto · 2026-09-18
- Astra for Law posts big legal benchmark gains as Mollick asks if labs will eat every AI vertical — emollick · 2026-09-18