Same thread: models that compute more per pass may no longer need to verbalize thinking

NeelNanda5 · x · 2026-09-11

Upper half of the same thread: the main way to detect misaligned models is reading their chain of thought, but if models do more per forward pass — as Astra's large no-CoT gains suggest — they don't need to verbalize their thinking, making them much harder to monitor.

Related event: Astra's No-CoT Reasoning Surge Raises Safety Concerns(7 posts)→

Original post →

More from Models

Models channel →