Do you need CoT to catch AI harm? Researchers clash over burden of proof

xeophon · x · 2026-09-04

Pushing back on claims that lacking chain-of-thought makes AI harder to monitor, S1r1u5 argues the burden of proof lies with claimants: show a concrete case where a model exploits a vulnerability via some sci-fi technique undetectable from its tool calls.

The debate responds to ZackKorman's post that you don't need reasoning tokens to detect real-world AI harm — you can just monitor the real world. It extends the ongoing discussion around Astra's CoT-monitorability.

Original post →

More from Safety

Safety channel →