ICML Paper: Reward Functions Are Only Partially Identifiable, Even With Infinite Data

gleech · x · 2026-09-07

gleech cites the ICML 2023 paper "Invariance in Policy Optimisation and Partial Identifiability in Reward Learning" (Skalse, Farrugia-Roberts, Russell, Abate, Gleave) as a pillar of the alignment thread: multiple reward functions often fit reward-learning data equally well, so the reward function is only partially identifiable even in the infinite-data limit. The paper formally characterizes partial identifiability across data sources like expert demonstrations and trajectory comparisons, and analyzes downstream impact on policy optimization — shifting the underlying true generator hurts objectives ("values") more than competence.

Original post →

More from Safety

Safety channel →