CoTs alone aren't sufficient, but removing explicit reasoning weakens defence-in-depth

schwarzjn_ · x · 2026-09-03

Responding in a debate on chain-of-thought observability, the author argues no one claims CoTs alone suffice — linear probes and targeted evals matter too — but the best practical strategy is defence-in-depth, and removing explicit reasoning eliminates an important, if imperfect, layer of it.

Related event: OpenAI's New Tech Reportedly Weakens CoT Monitorability, Sparking AI Safety Debate(32 posts)→

Original post →

More from AGI Musings

AGI Musings channel →