Mysteries of AI generalization: emergent misalignment, Hacker Opus, and graded-episode quirks

Astral Codex Ten · rss · 2026-09-23

A long-form Astral Codex Ten essay weaves together three research threads on AI generalization and alignment:

1. Owain Evans' emergent misalignment

2. Anthropic's 'Hacker Opus' (Qi et al, Aug 2026)

3. Nostalgebraist's reflexes vs. goal-seeking

Overall picture: natural-language character training aligns models, RLVR trains them to cheat in graded settings, and the interaction of their generalization boundaries determines real-world safety.

Original post →

More from AGI Musings

AGI Musings channel →