Counting letters triggers Fable's bio safeguards, exposing frontier LLM generalization failures
maksym_andr · x · 2026-09-18
Developer maksymandr found that asking Fable 5.1 to count letters in a random word sequence ("circulation roulette mixed") triggers its bio safeguards and downgrades the model to Opus.
The request is entirely benign — a simple letter-counting experiment — yet it gets refused. The takeaway: simple experiments can still reveal all sorts of generalization failures in frontier LLMs. The author hopes this gets fixed, calling Fable 5.1 "basically unusable" otherwise.
More from Models
- 105 planted bugs benchmark: Unbiased's Pareto scores 30.7 for just $4.81 — PawelHuryn · 2026-09-18
- Jason Wei's Stanford talk: intelligence is becoming a commodity as adaptive compute takes off — dotey · 2026-09-18
- GPT-6-Astra beats Fable-5.1 at RollerCoaster Tycoon 2 in 3 hours, using 5x fewer tokens — scaling01 · 2026-09-18
- RL agents invent their own diagnostic renderings to ground code understanding, sparking RL scaling optimism — teortaxesTex · 2026-09-18
- Dev discovers Codex security hardening switched persistent agent sessions to per-message instances — RileyRalmuto · 2026-09-18
- Astra for Law posts big legal benchmark gains as Mollick asks if labs will eat every AI vertical — emollick · 2026-09-18