前沿模型安全需监控潜伏空间,提防模型欺骗

belindazli · x · 2026-08-26

Safely deploying frontier models may require monitoring not only their behaviour, but also their latent space, especially in cases where models are scheming or deceptive. But in those cases, we must also evaluate how well we can trust these monitors --- a scheming model that could control its activations well could, in theory, confound probes that rely on those same activations. Some new work on measuring activation controllability in large language models.

原文链接 →

「安全」频道最新

更多「安全」频道 AI 资讯 →