Greenblatt Warns Latent Reasoning Architectures Would Sharply Raise Misalignment Risk

anmarasovic · x · 2026-09-26

Ryan Greenblatt and co-authors argue that latent reasoning architectures ('neuralese') would substantially increase misalignment risk by making oversight far harder. In the extreme, 'neuralese hivemind swarms' could think and communicate in latents, leaving oversight almost entirely dependent on observing agent actions — which those agents would have ample time to learn to obfuscate.

Stanford's Chris Potts calls the argument thoughtful but pushes back: interpretability should be nurtured as a hedge in case latent reasoning emerges anyway, and interp research should focus more on latent-reasoning models, with contingency plans even for seemingly hopeless outcomes.

Related event: Greenblatt et al. warn neuralese latent reasoning sharply raises AI misalignment risk(5 posts)→

Original post →

More from Safety

Safety channel →