Where AI Still Flunks IQ Tests: Spatial Reasoning, Visual Puzzles, and Complexity Limits
MIT Tech Review AI · rss · 2026-08-26
MIT Technology Review rounds up puzzles that have stumped AI models, inviting readers to try them while mapping where machine and human cognition diverge.
Key findings:
- Spatial reasoning: Despite visual inputs, language models fail badly at mental-rotation problems — they can't manipulate 3D objects the way spatial thinkers do.
- Memory as liability: When a puzzle resembles training data, models breeze past key differences and recite memorized answers. A 2024 Google/UIUC study on variants of Knights and Knaves confirmed this; on SimpleBench, humans spot the trick while top-tier models trip.
- Abstract & visual reasoning: On ARC-AGI, models that answer correctly often use byzantine, non-generalizable rules versus humans' simple visual concepts; encoding grids as number strings helps.
- Intuition inverted: Some problem suites exploit human knee-jerk biases that models — responding deliberatively — avoid.
- Complexity limits: Apple researchers found LLMs ace simple Tower of Hanoi and river-crossing puzzles but falter past six disks or people; the ZebraLogic study showed similar scaling limits on logic grids — though commentators debated whether this is a unique LLM flaw or just normal error accumulation.
Progress is fast: from Columbia's late-2024 finding that the best models solved only 18% of NYT Connections puzzles to near-perfect scores by early 2025 — but humans still hold distinct advantages.
More from Models
- Tiel-Coder-35B-A3B GGUF Released with Vision and Efficient Inference — peculiar-ragdoll · 2026-08-26
- Observation: Claude Starting With 'The Honest Answer...' Means It Failed — airesearch12 · 2026-08-26
- Trends of Most Used OpenRouter Models Over Time — Which-Breadfruit-926 · 2026-08-26
- AI models struggle with nuance in intelligence tests — nordicinst · 2026-08-26
- Devs Have Zero Model Loyalty: Ox Alpha Crushes DeepSeek V4 Flash Usage in 3 Days — FinanceYF5 · 2026-08-26
- OX Alpha tops OpenRouter usage by making the model free — Hesamation · 2026-08-26