Nat Lambert’s shorter lecture links over-optimization, reward hacking, and leaderboard gaming
natolambert · x · 2026-07-25
A shorter lecture from Nat Lambert on over-optimization, reward hacking, sycophancy, and verbosity.
- The talk covers the fundamentals and history of over-optimization through the lens of Goodhart’s Law.
- Lambert argues that rubrics, like reward models, are prone to their own kind of over-optimization, and that RLVR should be treated as a distinct phenomenon.
- The lecture also revisits signatures of misalignment, the limits of “just style” critiques, and the Llama 4 leaderboard-gaming episode.
- It is framed as a recap and reflection centered mainly on Chapter 14.
Related event: Lectures Explore LLM Over-Optimization and Reward Hacking(3 posts)→
More from AGI Musings
- AI companionship dissolves the friction real intimacy needs, warns long-form thread — YogeshMalik · 2026-09-11
- Why So Many AI Researchers Think the Machines Could Kill Everyone — wiredmagazine · 2026-09-11
- 'Hallucination' Is a Category Error: Naming AI 'Intelligence' Limits Our Imagination — Genaforvena · 2026-09-11
- Data engineering, not agent frameworks, is the real bottleneck for enterprise AI agents — dhruv2038 · 2026-09-11
- François Fleuret: Only Two Long-Term Futures — No Super AI, or Staying Fully Human With It — francoisfleuret · 2026-09-11
- IG reel debunking the 'winning the AI race against China' fallacy hits 500k likes — louisvarge · 2026-09-11