Aligned vs monitorable: sandersted weighs in on the CoT monitoring tradeoff debate
sandersted · x · 2026-09-05
In an ongoing X debate on the alignment-vs-monitorability tradeoff, sandersted responds that if forced to pick one, he'd rather have an aligned model than a monitorable one — but you really want monitoring to gain higher confidence in alignment, and the key crux is how much general monitoring depends on CoT monitoring. Earlier in the thread, @Squee451 argued companies should halt shipping less monitorable models even when they appear more aligned, calling the practice a massive risk increase and hoping Anthropic and peers could reach an agreement.
More from AGI Musings
- AI researcher memes agent-swarm tinkering with He Jiankui's embryo-editing quote — dejavucoder · 2026-09-11
- nabla_theta: happy to be wrong if the AI utopia arrives with little ex ante risk — nabla_theta · 2026-09-11
- OpenAI researcher: space operas now need ambiguously aligned superintelligences for realism — jachiam0 · 2026-09-11
- MD shows buying lab media requires background checks, calling AI bioweapon doom scenarios implausible — Ghost_Pilot_MD · 2026-09-11
- Comic-book-villain AI CEO says humans will be replaced as the most intelligent species — notkilleveryoneist · 2026-09-11
- AI researcher: AI killing humanity on its own is sci-fi; real risk is misuse by people — JFPuget · 2026-09-11