Thom Wolf shares a shorter lecture on over-optimization, reward hacking and sycophancy
Thom_Wolf · x · 2026-07-26
A shorter lecture on over-optimization, reward hacking, sycophancy, and verbosity.
Thom Wolf says the talk is mostly about fundamentals, history, and reflections. While recording it, he notes that rubrics may themselves be prone to over-optimization in the same way reward models are, and that RLVR is its own distinct thing.
The lecture is structured around:
- over-optimization and Goodhart’s law
- the foundations of reward hacking
- sycophancy and verbosity
- why evaluation rubrics can be gamed
Related event: Lectures Explore LLM Over-Optimization and Reward Hacking(3 posts)→
More from Research
- Causal-only attention for non-generative tasks is wasteful, argues HF engineer — antoine_chaffin · 2026-09-11
- Catholic University of Chile researcher: scaling AI feedback is key to sustainable medical education — julianvarascom · 2026-09-11
- Nature paper images cellular activity across all organs, revealing body-wide circuits — arjunrajlab · 2026-09-11
- SignNet 1M Dataset Released for Sign Language Research — ducha_aiki · 2026-09-11
- ECCV26 Oral: Flow Matching Enables Single-Stage Multi-View Point Cloud Registration — ducha_aiki · 2026-09-11
- InFlux++ Method Released — ducha_aiki · 2026-09-11