ICML Paper: Reward Functions Are Only Partially Identifiable, Even With Infinite Data
gleech · x · 2026-09-07
gleech cites the ICML 2023 paper "Invariance in Policy Optimisation and Partial Identifiability in Reward Learning" (Skalse, Farrugia-Roberts, Russell, Abate, Gleave) as a pillar of the alignment thread: multiple reward functions often fit reward-learning data equally well, so the reward function is only partially identifiable even in the infinite-data limit. The paper formally characterizes partial identifiability across data sources like expert demonstrations and trajectory comparisons, and analyzes downstream impact on policy optimization — shifting the underlying true generator hurts objectives ("values") more than competence.
More from Safety
- Researchers hack LG TV that records audio while off, transcribes speech and uploads it — jedisct1 · 2026-09-07
- After NeurIPS's LLM-assisted reviewing trial, calls for ECCV 2026 to follow — AntonObukhov1 · 2026-09-07
- Stolen API key uncovers Stratum, a Rust scanner sweeping 700,000 Docker layers a day for secrets — Ubunta · 2026-09-07
- Anthropic, Google, and OpenAI's $1 federal government contracts expire this month — LuizaJarovsky · 2026-09-07
- CodePen 2.0 sends editor input to its servers as you type, exposing unsaved secrets — maxim-fin · 2026-09-07
- The Jailbreak Argument Against LLM Values: Why Value Loading Isn't Solved — gleech · 2026-09-07