Latent Reasoning ('Neuralese') Would Sharply Raise Misalignment Risk, Greenblatt Argues

HaydnBelfield · x · 2026-09-25

Ryan Greenblatt and colleagues argue that latent reasoning architectures ('neuralese') would substantially increase misalignment risk by making oversight much harder:

The argument was endorsed and amplified by ancadianadragan and Haydn Belfield.

Related event: Researchers Warn Latent Reasoning ('Neuralese') Sharply Raises AI Misalignment Risk(5 posts)→

Original post →

More from Safety

Safety channel →