Was Claude's gibberish fixed by capping KL divergence in RL? One theory

burny_tech · x · 2026-10-11

burnytech speculates that Opus 5.5 speaks in incomprehensible language far less than previous Claude models because Anthropic capped KL divergence more tightly during RL training. KL constraints limit how far the policy can drift from the base model, and tighter caps could suppress the tendency to degrade into unreadable output. This is an unverified personal theory, but it touches on a real technical question about how KL regularization shapes language degeneration in RLHF.

Original post →

More from Models

Models channel →