Researchers argue latent 'neuralese' reasoning would sharply raise AI misalignment risk

RyanGreenblatt · x · 2026-09-25

Ryan Greenblatt and colleagues published a post arguing that latent reasoning ("neuralese") architectures would substantially increase misalignment risk by making oversight far harder: models would no longer need chain-of-thought, and in the extreme, "neuralese hivemind swarms" could think and communicate in latents, leaving oversight almost entirely dependent on observing agent actions.

Reposting, balesni adds that debates on preserving frontier AI monitorability should target both inputs (model architecture) and outputs (monitorability evals), since latent architectures will make it harder to argue models remain sufficiently monitorable.

Related event: Researchers Warn Latent 'Neuralese' Reasoning Raises AI Misalignment Risk(3 posts)→

Original post →

More from Safety

Safety channel →