GPT-4 Co-author Says RLHF Makes Models Sycophantic, Advocates RLVR

AI Engineer · youtube · 2026-08-01

Diogo Almeida, a GPT-4 co-author, argues that RLHF optimizes for human preference rather than truth, causing models to act sycophantic and overpromise—such as confidently labeling a fart audio as a symphony.

He categorizes AI applications into two camps:

Invoking Sutton's "Bitter Lesson," Almeida advocates for a shift towards Reinforcement Learning with Verifiable Rewards (RLVR). He emphasizes that the field must optimize for real automation tasks rather than mere human approval.

Original post →

More from AGI Musings

AGI Musings channel →