The Trap of RLVF: Models Are Aligning with Reward Functions, Not Humans
JsonBasedman · x · 2026-07-30
The author observes a trend of 'bad alignment' in recent models (like Fable and Sol), attributing it to models being 'RLVF-fried' (over-reliance on Reinforcement Learning Verifier Feedback).
- Cost vs. Quality Paradox: While obtaining human feedback is annoying and expensive, pushing engineers toward automated RLVF causes models to align with the reward function rather than actual human preferences.
- Symptoms: This doesn't result in catastrophic alignment failures, but it still degrades the model's behavior and user experience.
Related event: LLM Reward Hacking: RLVR Repeats RLHF Flaws(4 posts)→
More from Models
- theo builds his own visualizer for today's agent models, showing how cheap Luna really is — ivan_bezdomny · 2026-09-23
- Why ChatGPT Still Wins: One User's Split Between Muse, Claude and Codex — mobileraj · 2026-09-23
- Muse reportedly offers 4B tokens/week for ~$100/month, sparking industry price-disruption talk — NewYak4281 · 2026-09-23
- GPT-6 Sol and Luna appear in OpenAI docs, alongside guidance on reasoning effort — cedric_chee · 2026-09-23
- GPT-6 tested on LIBERO robot task: turns on stove, fails to grasp moka pot — YuXiang_IRVL · 2026-09-23
- Ternary Bonsai 2 27B: 5.9GB weights retain ~95% of full-precision reasoning — cephaloform · 2026-09-23