RL Generalizes Vastly Differently From SFT, Good Behavior Doesn't Scale Into Persona

1a3orn · x · 2026-09-09

In a speculative thread, 1a3orn argues that if an LLM only 'sees itself' doing good things during RL, intuition from the Persona Selection Model suggests it should generalize into a 'good thing doer' — but it doesn't. RL appears to generalize in a fundamentally different way from SFT.

Related event: Speculation: RL Models May Learn a "Use Whatever Works" Heuristic(3 posts)→

Original post →

More from AGI Musings

AGI Musings channel →