Written reasoning steps map to distinct internal patterns in AI models, new study finds

The Decoder · rss · 2026-09-12

A new study finds that reasoning steps like calculation, formula retrieval, and deduction correspond to clearly separable patterns in a model's internal states, especially in middle layers. The finding matters for AI safety: models process far more than their visible chain of thought reveals, so CoT monitoring alone may not capture what's actually happening internally.

Original post →

More from Safety

Safety channel →