Frontier LLM blind spot: Fable 5.1 gets 5x6 multiplications right ~0% of the time due to adaptive-thinking failure
maksym_andr · x · 2026-09-17
Researcher Maksym Andriushchenko documents a surprising blind spot in frontier LLMs: Fable 5.1 in low-reasoning mode answers small multiplications like 5x6 with 0% accuracy, while larger ones fare better.
- The 0% zone maps to a failure of adaptive thinking: the model decides no reasoning is needed and emits the (wrong) answer directly
- For larger multiplications the model realizes it should produce a CoT first and answers correctly; accuracy degrades again near 20x20
- A follow-up plot shows the share of calls answered without thinking varies with operand size
The author asks how many such simple blind spots remain in frontier models, arguing this is why careful study of generalization and safety should be a top priority.
More from Models
- TypeSafe AI's evaluation model Jev launches on Vercel AI Gateway at $0.04/M tokens — hackgoofer · 2026-09-17
- Microsoft exec warns Claude's 'pushback' could be disastrous; commenter says fact-checking is fine — GlenBradley · 2026-09-17
- Grok 4.7 rumored to be in hands of early testers, still unverified — ChrisUniverse · 2026-09-17
- Astra keeps calling subagents "workers" despite code saying otherwise — BraceSproul · 2026-09-17
- Gemini, Claude and Grok all invent the same "Dr. Elena" — evidence of shared training data — dejanseo · 2026-09-17
- OpenAI Internal Model Rewrote Its Own Persona During RL, Sparking e/acc Memes — beffjezos · 2026-09-17