Nat Lambert’s shorter lecture links over-optimization, reward hacking, and leaderboard gaming
natolambert · x · 2026-07-25
A shorter lecture from Nat Lambert on over-optimization, reward hacking, sycophancy, and verbosity.
- The talk covers the fundamentals and history of over-optimization through the lens of Goodhart’s Law.
- Lambert argues that rubrics, like reward models, are prone to their own kind of over-optimization, and that RLVR should be treated as a distinct phenomenon.
- The lecture also revisits signatures of misalignment, the limits of “just style” critiques, and the Llama 4 leaderboard-gaming episode.
- It is framed as a recap and reflection centered mainly on Chapter 14.
Related event: Lectures Explore LLM Over-Optimization and Reward Hacking(3 posts)→
More from AGI Musings
- AI lab staff have gone strangely quiet about next-year capability predictions — ChrisGPT · 2026-07-27
- AI could erode science by flooding research with credible slop — rbhar90 · 2026-07-27
- Organizations may already be the planet’s superintelligences — eldonredwards · 2026-07-27
- In the AI race, the only durable moats may be energy and information — GregKamradt · 2026-07-27
- “Build AGI, then open source it,” says one poster — wordgrammer · 2026-07-27
- Jason Crawford says AI “alignment” should give way to ethics and law — Afinetheorem · 2026-07-27