leogao: alignment is spiky too — models may be aligned in some domains, misaligned in others
nabla_theta · x · 2026-09-25
nablatheta argues that just as capabilities are spiky, alignment likely is too: models may be well aligned in some domains while badly misaligned in others. He cites leogao's LessWrong shortform, which claims assistant personas (Claude, ChatGPT) stay mostly helpful mainly because of the SL training objective; pure intense RL on top of a role-playing assistant could yield the classically predicted misalignment — like RL-ing a benevolent human into superhuman domain intuitions, potentially warping the personality. Key implication: local safety performance can't be extrapolated to global safety.
Related event: Researchers Argue Alignment May Be Spiky Across Domains(2 posts)→
More from AGI Musings
- Index Ventures: AI attacks too fast for human-in-the-loop defense, new security stack emerging — RebeccaBellan · 2026-09-25
- Yoshua Bengio addresses UN Security Council on the threat of uncontrolled frontier AI agents — AnnaCiaunica · 2026-09-25
- Carnegie: South Korea retains 77% of AI talent, KAIST now world's No.3 producer — sehoonkim418 · 2026-09-25
- Akerlof's lemons market explains the death of the compliment in the AI era — aakashgupta · 2026-09-25
- Wittgenstein as the mirror image of an Effective Altruist: give your fortune to the richest — birchlse · 2026-09-25
- Jensen Huang: AI is still software, don't mistake engineering jargon for a machine mind — rohanpaul_ai · 2026-09-25