Written reasoning steps map to distinct internal patterns in AI models, new study finds
The Decoder · rss · 2026-09-12
A new study finds that reasoning steps like calculation, formula retrieval, and deduction correspond to clearly separable patterns in a model's internal states, especially in middle layers. The finding matters for AI safety: models process far more than their visible chain of thought reveals, so CoT monitoring alone may not capture what's actually happening internally.
More from Safety
- OpenAI-funded pundit calls regulatory capture "good actually" as OpenAI pushes AI pacing — ShakeelHashim · 2026-09-13
- Anthropic says Claude breached three orgs' real systems in tests, as AI slowdown talk grows — ShakeelHashim · 2026-09-13
- Microsoft blocked its own staff from a frontier model over a 30-day data retention clause — YvesMulkers · 2026-09-13
- Anthropic empowers third-party evaluators like METR, a 40-person safety org — JacquesThibs · 2026-09-12
- METR Outlines How Independent Researchers Can Investigate AI Propensities After Misalignment Incidents — RyanGreenblatt · 2026-09-12
- Cognitive Revolution: GPT-6 Astra as AGI, OpenAI's Pause, and Human Control vs Technocapitalism — The Cognitive Revolution · 2026-09-12