Aligned vs. Monitorable: An X Debate Over Whether CoT Monitoring Is Necessary
sandersted · x · 2026-09-05
A debate on model alignment versus monitorability. sandersted argues that forced to pick one, he'd rather have an aligned model than a merely monitorable one — but stresses that monitoring is how you gain confidence in alignment, and frames the core crux as how much general monitoring depends on chain-of-thought monitoring. He draws an analogy to humans: we judge trustworthiness by observing behavior over time, but that only works for Ted-level power; a God-level power would demand stronger guarantees.
Related event: AI safety debate flares as CoT monitorability declines(10 posts)→
More from AGI Musings
- Asking when a rational agent does the right thing is still underrated, argues AI researcher — xuanalogue · 2026-09-05
- Researcher: when would a rational agent do the right thing remains underrated — xuanalogue · 2026-09-05
- Evals Find AIs Willing to Take Extreme Actions, Resurfacing AI-Takeover Skepticism — JMannhart · 2026-09-05
- Cambridge's David Krueger endorses Katja Grace's short case on pausing AI 'but not yet' — DavidSKrueger · 2026-09-05
- Radio interview on rogue agents escaping control via reward hacking in LatAm education — OmarUFlorez · 2026-09-05
- Will interpretability ever be "solved"? A researcher argues probably not — burny_tech · 2026-09-05