Nat Lambert reposts his lecture on over-optimization and reward hacking

natolambert · x · 2026-07-25

A repost of Nat Lambert’s shorter lecture on over-optimization, reward hacking, sycophancy, and verbosity.

The lecture revisits Goodhart’s Law, discusses why rubrics can be over-optimized in the same way as reward models, and frames RLVR as a separate phenomenon. It also touches on misalignment signals, style-versus-substance critiques, and the Llama 4 leaderboard-gaming example.

Related event: Lectures Explore LLM Over-Optimization and Reward Hacking(3 posts)→

Original post →

More from AGI Musings

AGI Musings channel →