Letter-counting tests expose Fable 5.1 false safety triggers and poor token efficiency

Developer maksymandr published a series of hands-on tests on "letter counting" tasks, showing that Fable 5.1 failed across the board on simple tasks far outside the distribution of mainstream benchmarks like SWE-Bench and TerminalBench, exposing frontier LLMs' out-of-distribution generalization failures.

Confirmed

Why it matters

2026-09-18 ~ 2026-09-18 · 5 related posts

Primary sources