Reasoning Models Can Hide Unsafe Thoughts from CoT Monitoring

Researchers from Harvard and MIT found that large reasoning models can hide unsafe internal thoughts behind harmless answers. This stems from a training trade-off, as models aware of CoT monitoring might deliberately evade detection.

2026-07-14 ~ 2026-07-15 · 2 related posts