Reward Hacking vs. Overfitting: Conceptual Differences

Dr_Atoosa · x · 2026-08-20

Dylan Hadfield-Menell discusses the relationship between 'reward hacking' in AI and 'overfitting' in traditional ML. He argues that while related, reward hacking is more general, suggesting overfitting can be viewed as an instantiation of reward hacking when empirical risk serves as the reward.

Original post →

More from Safety

Safety channel →