Exploring Claude's Weird Outputs: RL Training Might Cause OOD

cephaloform · x · 2026-07-30

Addressing the anomalous behavior of Claude's base model generating user-modeling-like text, a developer offers a technical explanation. They suggest the weirdness might stem from the Reinforcement Learning (RL) training phase.

Specifically, masking user replies during RL loss computation causes the model to rely on base user modeling from mid-training during inference. The RL delta could be pulling the model slightly out of distribution (OOD), resulting in these confusing generations.

Related event: Developers Attribute Claude's Odd Outputs to User Modeling, Not Thought(3 posts)→

Original post →

More from Models

Models channel →