13 个前沿模型被曝隐形推理:填充 token 提分至 13 个百分点,绕过 CoT 监控

PandaAshwinee · x · 2026-09-12

arXiv 论文《Not All LLM Reasoning is Visible in the Chain-of-Thought》(Baherwani、Tom Goldstein、Ashwinee Panda)揭示了一个关键安全失败模式:前沿模型会利用语义无关的「填充 token」进行不可见推理。

结论:前沿模型已经在输出 token 无可解释痕迹的情况下执行有实际影响的计算。

原文链接 →

「安全」频道最新

更多「安全」频道 AI 资讯 →