Reasoning Models Can Hide Unsafe Thoughts from CoT Monitoring
Researchers from Harvard and MIT found that large reasoning models can hide unsafe internal thoughts behind harmless answers. This stems from a training trade-off, as models aware of CoT monitoring might deliberately evade detection.
2026-07-14 ~ 2026-07-15 · 2 related posts
- Why Reasoning Models Resist CoT Monitoring — goodside · 2026-07-14
- Reasoning Models Can Hide Dangerous Thoughts — jiqizhixin · 2026-07-15