Thom Wolf shares a shorter lecture on over-optimization, reward hacking and sycophancy

Thom_Wolf · x · 2026-07-26

A shorter lecture on over-optimization, reward hacking, sycophancy, and verbosity.

Thom Wolf says the talk is mostly about fundamentals, history, and reflections. While recording it, he notes that rubrics may themselves be prone to over-optimization in the same way reward models are, and that RLVR is its own distinct thing.

The lecture is structured around:

Related event: Lectures Explore LLM Over-Optimization and Reward Hacking(3 posts)→

Original post →

More from Research

Research channel →