A 1978 logic puzzle book is catching chatbots that memorize instead of reason
JafarNajafov · x · 2026-09-25
A long thread traces Raymond Smullyan, his 1978 puzzle book What Is the Name of This Book?, and how its knight-and-knave riddles became a tool for exposing chatbots that fake reasoning.
- The man: hooked on logic at six by an April Fool's riddle, dropped out of high school, performed card tricks as "Five-Ace Merrill," and earned a Princeton PhD under Alonzo Church (Turing's advisor).
- The puzzles: knights always tell truth, knaves always lie; each puzzle forces you to hold two possible worlds and watch one collapse — 200+ of them, getting meaner.
- The 2024 twist: AI researchers used these riddles to test whether models reason or memorize. Models aced the training puzzles, then collapsed when researchers swapped who was the knight or changed a single statement.
- Takeaway: most benchmarks reward right answers; Smullyan's puzzles punish the moment you stop thinking and start pattern matching — catching a flaw billion-dollar benchmarks missed.
More from Models
- OpenAI reportedly prepping GPT-6 Cyber security model for DevDay — emmanuelvivier · 2026-09-25
- Anthropic Publishes Official Prompting Guide for Claude Opus 5.5: Calibrate Effort, Don't Max It — CodeByPoonam · 2026-09-25
- OrcaSAQ-2-27B Trends on Hugging Face: 3-Bit Mixed-Precision Qwen3-Based Reasoning Model — orcarouter · 2026-09-25
- AVIS paper accepted at NeurIPS 2026: jointly scaling visual context and reasoning per query — CSProfKGD · 2026-09-25
- "I feel the AGI": developer says Opus 5.5 cracked video editing, and he's the bottleneck — LeviTurk · 2026-09-25
- Not every task needs frontier models: local Qwen 4 27B is pulling users away — haider1 · 2026-09-25