Local Bad Behavior In RL Doesn't Produce Emergent Misalignment, Per Hacker-Opus

1a3orn · x · 2026-09-09

Continuing his thread, 1a3orn cites the hacker-opus observation that when an LLM 'sees itself' doing bad things during RL, we currently don't get emergent misalignment: bad behavior stays local and fails to generalize into globally bad behavior.

Related event: Local RL Behavior Doesn't Generalize into Global Traits(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →