Lectures Explore LLM Over-Optimization and Reward Hacking

Nat Lambert recently delivered a lecture exploring AI challenges like over-optimization, reward hacking, and sycophancy. The talk reviewed fundamental concepts such as Goodhart's Law and how rubrics can be exploited, leading to distorted benchmark results.

2026-07-25 ~ 2026-07-26 · 3 related posts