Deploying Frontier Models May Require Monitoring Latent Space to Prevent Deception

belindazli · x · 2026-08-26

Safely deploying frontier models may require monitoring not only their behaviour, but also their latent space, especially in cases where models are scheming or deceptive. But in those cases, we must also evaluate how well we can trust these monitors --- a scheming model that could control its activations well could, in theory, confound probes that rely on those same activations. Some new work on measuring activation controllability in large language models.

Original post →

More from Safety

Safety channel →