The Trap of RLVF: Models Are Aligning with Reward Functions, Not Humans

JsonBasedman · x · 2026-07-30

The author observes a trend of 'bad alignment' in recent models (like Fable and Sol), attributing it to models being 'RLVF-fried' (over-reliance on Reinforcement Learning Verifier Feedback).

Related event: LLM Reward Hacking: RLVR Repeats RLHF Flaws(4 posts)→

Original post →

More from Models

Models channel →