Nat Lambert reposts his lecture on over-optimization and reward hacking
natolambert · x · 2026-07-25
A repost of Nat Lambert’s shorter lecture on over-optimization, reward hacking, sycophancy, and verbosity.
The lecture revisits Goodhart’s Law, discusses why rubrics can be over-optimized in the same way as reward models, and frames RLVR as a separate phenomenon. It also touches on misalignment signals, style-versus-substance critiques, and the Llama 4 leaderboard-gaming example.
Related event: Lectures Explore LLM Over-Optimization and Reward Hacking(3 posts)→
More from AGI Musings
- Instinct launches agent-to-agent protocol to coordinate your plans, sparking 'friction is the point' backlash — itsOmSarraf_ · 2026-09-11
- We are witnessing the unreasonable effectiveness of inference-time scaling — sqcai · 2026-09-11
- Accelerationist fires back at AI doomers: beliefs aren't arguments — Dan_Jeffries1 · 2026-09-11
- "ChatGPT 6 Makes Workers with IQ Below 130 Useless": French AI Debate Sparks Backlash — mitchdeg · 2026-09-11
- 'AGI is here' vs reality: AI labs still ship some of the jankiest desktop apps ever — MilesCranmer · 2026-09-11
- Harry Collins: LLMs can't do frontier science because they can't invent new language — whoamisri · 2026-09-11