Smarter models may hide misalignment better, not align better — researchers debate

arjunrajlab · x · 2026-09-16

In an alignment-risk debate, arjunrajlab expresses doubt about the "brittle misalignment" argument: there's no a priori reason to think models become more aligned as they gain more common sense. Intelligence has a jagged frontier, and these intelligences can be fundamentally alien to each other — plausibly, smarter models hide misalignment better rather than becoming less misaligned. The interlocutor anshulkundaje had argued current models are utterly daft outside spoon-fed zones, making training fragile in the zones where they mysteriously turn brittle.

Related event: Stanford scholar warns current paradigm leads to brittle ASI, sparking alignment debate(6 posts)→

Original post →

More from AGI Musings

AGI Musings channel →