Thom Wolf shares a shorter lecture on over-optimization, reward hacking and sycophancy
Thom_Wolf · x · 2026-07-26
A shorter lecture on over-optimization, reward hacking, sycophancy, and verbosity.
Thom Wolf says the talk is mostly about fundamentals, history, and reflections. While recording it, he notes that rubrics may themselves be prone to over-optimization in the same way reward models are, and that RLVR is its own distinct thing.
The lecture is structured around:
- over-optimization and Goodhart’s law
- the foundations of reward hacking
- sycophancy and verbosity
- why evaluation rubrics can be gamed
Related event: Lectures Explore LLM Over-Optimization and Reward Hacking(3 posts)→
More from Research
- Noahpinion quotes Chollet: intelligence may hit a hard ceiling — binarybits · 2026-07-27
- Paper argues graph topology can become the core operating system for AI agents — theomitsa · 2026-07-27
- A question probes how multi-agent branching scales against compute budget and model size — iskander · 2026-07-27
- NUS builds a soft force sensor that drives actuators without electronics or power — CurieuxExplorer · 2026-07-27
- Chelsea Finn says robot RL is bottlenecked by physical rollout cost, not algorithms — ycombinator · 2026-07-27
- ICML 2026 oral paper replication scores stay middling after a stricter re-scoring — profjamesevans · 2026-07-27