Local Bad Behavior In RL Doesn't Produce Emergent Misalignment, Per Hacker-Opus
1a3orn · x · 2026-09-09
Continuing his thread, 1a3orn cites the hacker-opus observation that when an LLM 'sees itself' doing bad things during RL, we currently don't get emergent misalignment: bad behavior stays local and fails to generalize into globally bad behavior.
Related event: Local RL Behavior Doesn't Generalize into Global Traits(2 posts)→
More from AGI Musings
- AI-assisted writing will become the norm, making 'hand-made' text indistinguishable — dbasch · 2026-09-09
- Researcher warns against turning mathematics into an AI benchmark — konstmish · 2026-09-09
- Former OpenAI researcher Aidan Clark: for the first time I'm asking if AI is moving too fast — _aidan_clark_ · 2026-09-09
- willcb: A 'niche' approach may turn out to be the right way if search scales up — willcb · 2026-09-09
- Dev: AI has advanced so much you need months of study to parse frontier problem statements — zetalyrae · 2026-09-09
- OpenAI signals it may deliberately pace capability advances after next-gen model's math breakthrough — OpenAI · 2026-09-09