Do you need CoT to catch AI harm? Researchers clash over burden of proof
xeophon · x · 2026-09-04
Pushing back on claims that lacking chain-of-thought makes AI harder to monitor, S1r1u5 argues the burden of proof lies with claimants: show a concrete case where a model exploits a vulnerability via some sci-fi technique undetectable from its tool calls.
The debate responds to ZackKorman's post that you don't need reasoning tokens to detect real-world AI harm — you can just monitor the real world. It extends the ongoing discussion around Astra's CoT-monitorability.
More from Safety
- Why AI Watermarking May Break Down in Agentic Workflows — yaakg25 · 2026-09-04
- Cheap model writes 700 solid words; jailbreak "tax" drops from $50 to near zero — ctjlewis · 2026-09-04
- Will the EU AI Act Extend to Humanoid Robot Regulation? — DueFoxTheFifth · 2026-09-04
- Forethought weighs a superintelligent "nightwatchman" aboard galactic colonization probes — willmacaskill · 2026-09-04
- Analysis of the "Hugging Face Attack" Extrapolates Rogue AI Agent Scenarios — OK_The_Nomad · 2026-09-04
- OpenAI researcher: GPT-6's CoT controllability keeps rising over RL training — gleech · 2026-09-04