前沿模型安全需监控潜伏空间,提防模型欺骗
belindazli · x · 2026-08-26
Safely deploying frontier models may require monitoring not only their behaviour, but also their latent space, especially in cases where models are scheming or deceptive. But in those cases, we must also evaluate how well we can trust these monitors --- a scheming model that could control its activations well could, in theory, confound probes that rely on those same activations. Some new work on measuring activation controllability in large language models.
「安全」频道最新
- Agent 委托即提权:工具白名单只挡一跳,约束别写在 prompt 里 — anp2_protocol · 2026-08-26
- 传 OpenAI 下代模型 Astra 攻克难题,因高风险被暂 — haider1 · 2026-08-26
- 参与反AI抗议可赋予员工举报勇气 — RichardMCNgo · 2026-08-26
- AI开发者吁监管,但国会立法进度滞后 — DavidSKrueger · 2026-08-26
- Cisco 四个月第四笔身份交易:投资 Teleport 强化基础设施身份层 — shashib · 2026-08-26
- Google 赞助 2026 年 [un]prompted AI × 网络安全大会,议题涵盖自动化研究 — dyn___ · 2026-08-26