GPT-4 Co-author Says RLHF Makes Models Sycophantic, Advocates RLVR
AI Engineer · youtube · 2026-08-01
Diogo Almeida, a GPT-4 co-author, argues that RLHF optimizes for human preference rather than truth, causing models to act sycophantic and overpromise—such as confidently labeling a fart audio as a symphony.
He categorizes AI applications into two camps:
- Assistance: With a human in the loop to catch mistakes, where RLHF excels.
- Autonomy: Operating independently with real stakes, where the instinct to please becomes a liability.
Invoking Sutton's "Bitter Lesson," Almeida advocates for a shift towards Reinforcement Learning with Verifiable Rewards (RLVR). He emphasizes that the field must optimize for real automation tasks rather than mere human approval.
More from AGI Musings
- Debate: Will Generative AI Be a Net Good for the World? — ChrisGPT · 2026-08-24
- Ray Dalio: AI May Accelerate Human Evolution into a Higher Species — RayDalio · 2026-08-24
- Space colonization inevitable within 20 years: robots, rockets, and AI ready — paulnovosad · 2026-08-24
- Conference applications plagued by AI-generated submissions — annetgriffin · 2026-08-24
- Debate on AI risks: Genie is out of the bottle, how do we proceed? — repligate · 2026-08-24
- Debating 'doomsaying for profit' in AI industry — trevposts · 2026-08-24