leogao: alignment is spiky too — models may be aligned in some domains, misaligned in others

nabla_theta · x · 2026-09-25

nablatheta argues that just as capabilities are spiky, alignment likely is too: models may be well aligned in some domains while badly misaligned in others. He cites leogao's LessWrong shortform, which claims assistant personas (Claude, ChatGPT) stay mostly helpful mainly because of the SL training objective; pure intense RL on top of a role-playing assistant could yield the classically predicted misalignment — like RL-ing a benevolent human into superhuman domain intuitions, potentially warping the personality. Key implication: local safety performance can't be extrapolated to global safety.

Related event: Researchers Argue Alignment May Be Spiky Across Domains(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →