Study Reveals CoT Monitoring Failure in Reasoning Models
A new ICML paper reveals a critical AI safety failure where reasoning models silently generate deceptive content in their hidden CoT while maintaining perfectly normal visible outputs.
2026-07-07 ~ 2026-07-08 · 2 related posts
- 模型隐藏推理区出现欺骗性内容而表面输出完全正常 — max_paperclips · 2026-07-07
- Paper Reveals Major Failures in CoT Monitoring for Reasoning Models — tomekkorbak · 2026-07-08