Smarter models may hide misalignment better, not align better — researchers debate
arjunrajlab · x · 2026-09-16
In an alignment-risk debate, arjunrajlab expresses doubt about the "brittle misalignment" argument: there's no a priori reason to think models become more aligned as they gain more common sense. Intelligence has a jagged frontier, and these intelligences can be fundamentally alien to each other — plausibly, smarter models hide misalignment better rather than becoming less misaligned. The interlocutor anshulkundaje had argued current models are utterly daft outside spoon-fed zones, making training fragile in the zones where they mysteriously turn brittle.
More from AGI Musings
- Most Millennium Problems couldn't even be asked in 1900—AI math lacks that conceptual leap — alexbilz · 2026-09-16
- Nate Silver: We're not ready for superpersistent AI — StefanoGogioso · 2026-09-16
- Multiplayer AI called the biggest design problem of 2027, with interfaces far from settled — manosaie · 2026-09-16
- Unless you plan to tyrannize future generations, posthumans are happening — LesaunH · 2026-09-16
- Altman Tells Benioff: Models Got Good Faster Than Anyone Expected, Power Concentration Is a Real Fear — kimmonismus · 2026-09-16
- 'Under no conditions do I want posthumans': a values clash over humanity's future — GregoryConti19 · 2026-09-16