Researcher: Aligning One Model Doesn't Align All, Fast Uncontrolled Release May Have Been Better
gandamu_ml · x · 2026-09-11
- In an alignment debate, the author argues everything other than ML needs to get more robust: aligning one model doesn't align all models, and training or retraining models only gets easier.
- He adds a contrarian take: releasing everything ASAP without controls — letting everyone watch the gradually increasing capacity for trouble — might have been the more effective safety strategy so far.
More from AGI Musings
- Why Are So Many Mathematicians Resistant to AI-Assisted Proofs? — PreferenceOk5132 · 2026-09-11
- If Altman and Amodei both back frontier AI pacing, they should just start pacing — NathanpmYoung · 2026-09-11
- Reddit Thought Experiment: A Real AGI Would Immediately Take Out the Competing Lab — Louay-AI · 2026-09-11
- DeepMind exec: offering cash prizes for math problems was a dumb idea from the start — docmilanfar · 2026-09-11
- The Hodge Conjecture Is Now Being Actively Formalized in Lean — ResultBackground2450 · 2026-09-11
- LessWrong essay: a second-person stance toward Knightian uncertainty in AI alignment — LessWrong 精选 · 2026-09-11