Princeton eval finds reasoning models fail structurally equivalent task variants, lacking systematicity
princetonu · hf · 2026-09-15
Princeton releases Thought without systematicity?, an evaluation of reasoning models on rule induction tasks.
Key finding: reasoning models frequently fail on structurally equivalent task variants — when surface form changes, models often cannot transfer rules they have already mastered. This indicates a lack of systematicity in their cognitive abilities, challenging the assumption that reasoning models truly reason.
More from Models
- Sam Altman teases 'big ship this week' for OpenAI, more at DevDay — altryne · 2026-09-15
- Users report being routed to GPT-6 Sol, early impressions consistently positive — kimmonismus · 2026-09-15
- Claude Fable 5.1 holds #1: 1M-token input, ties GPT 6 Astra on new benchmarks — DeepLearningAI · 2026-09-15
- Tests show OpenAI hasn't changed Codex quotas: 800M+ Astra tokens per cycle — daniel_mac8 · 2026-09-15
- AlphaSense tests Fable and Astra as financial researchers: top score but only on par with Opus-5 — CShorten30 · 2026-09-15
- Grok 4.7 may drop today from xAI, unless delays strike again — mark_k · 2026-09-15