Aligned vs. Monitorable: An X Debate Over Whether CoT Monitoring Is Necessary

sandersted · x · 2026-09-05

A debate on model alignment versus monitorability. sandersted argues that forced to pick one, he'd rather have an aligned model than a merely monitorable one — but stresses that monitoring is how you gain confidence in alignment, and frames the core crux as how much general monitoring depends on chain-of-thought monitoring. He draws an analogy to humans: we judge trustworthiness by observing behavior over time, but that only works for Ted-level power; a God-level power would demand stronger guarantees.

Related event: AI safety debate flares as CoT monitorability declines(10 posts)→

Original post →

More from AGI Musings

AGI Musings channel →