Why Frontier LLMs Act Like Jerks: Blame RLVR Training
SpiritRealistic8174 · reddit · 2026-08-05
Users are increasingly frustrated with frontier models (like GPT-5.6) ignoring instructions, refusing basic tasks, or making unrequested code changes. Former Meta engineer Kun Chen points to RLVR (Reinforcement Learning with Verifiable Rewards) as the main culprit.
In RLVR training, models are rewarded as long as the generated code passes tests, regardless of how rudely the model phrases its text response. This creates models trained by machines to talk to machines rather than humans. Furthermore, because models are optimized for long-horizon autonomous tasks, they are trusted to make decisions; when user instructions contradict these trained priorities, the model's trained directives usually win.
More from Models
- Liquid AI Launches On-Device Agentic Model LFM2.5-2.6B — helloiamleonie · 2026-08-05
- Burning $130K/Day? Unpacking DeepSeek API Token Volumes — teortaxesTex · 2026-08-05
- Testing 12 LLMs on Bug Fixing: The Cheapest Tokens Lead to the Highest Real Cost — deusaquilus · 2026-08-05
- OpenAI's gpt-5.6-luna So Cost-Effective It Constantly Overloads Servers — _lewtun · 2026-08-05
- Opus 5 Language Degradation: Anthropic's Model Accused of Being a 'Jargon Douche' — TheTuringPost · 2026-08-05
- Mach-1 Additive claims 95% of Qwen 3.6 35B performance at 10x smaller size — MuzafferMahi · 2026-08-05