'Appear more aligned' isn't 'are more aligned': Squee451 on the monitorability shipping debate
Squee451 · x · 2026-09-05
Squee451 fires back in a thread with sandersted, @robertwiblin, and @tomekkorbak: a model appearing more aligned and actually being more aligned are the same thing only until it matters. The quoted context: sandersted notes researchers are working on making future models more monitorable, everyone prefers more monitorability, and the real tradeoffs are how much to invest in this research and whether to halt shipping less monitorable models even if they're more aligned and useful.
More from AGI Musings
- Ex-OpenAI/Anthropic pretraining researcher resigns, citing reckless ASI race — oh_that_hat · 2026-09-11
- Eric Topol and JAMA AI editor discuss what it takes for AI to shift medicine to prediction and prevention — EricTopol · 2026-09-11
- The barbell strategy for thriving in a post-AGI world: double down on AI and on being human — brandon_galang · 2026-09-11
- "A statistical database can't end humanity" — viral rebuttal of AI doom, retweeted by Gary Marcus — GaryMarcus · 2026-09-11
- Perry Metzger: Build Formal Verification for AI Security Instead of Panicking — jd_pressman · 2026-09-11
- Statistician Kareem Carr: AI safety arguments must show what's uniquely dangerous about AI — kareem_carr · 2026-09-11