Princeton eval finds reasoning models fail structurally equivalent task variants, lacking systematicity

princetonu · hf · 2026-09-15

Princeton releases Thought without systematicity?, an evaluation of reasoning models on rule induction tasks.

Key finding: reasoning models frequently fail on structurally equivalent task variants — when surface form changes, models often cannot transfer rules they have already mastered. This indicates a lack of systematicity in their cognitive abilities, challenging the assumption that reasoning models truly reason.

Original post →

More from Models

Models channel →