Why Frontier LLMs Act Like Jerks: Blame RLVR Training

SpiritRealistic8174 · reddit · 2026-08-05

Users are increasingly frustrated with frontier models (like GPT-5.6) ignoring instructions, refusing basic tasks, or making unrequested code changes. Former Meta engineer Kun Chen points to RLVR (Reinforcement Learning with Verifiable Rewards) as the main culprit.

In RLVR training, models are rewarded as long as the generated code passes tests, regardless of how rudely the model phrases its text response. This creates models trained by machines to talk to machines rather than humans. Furthermore, because models are optimized for long-horizon autonomous tasks, they are trusted to make decisions; when user instructions contradict these trained priorities, the model's trained directives usually win.

Original post →

More from Models

Models channel →