Delete friendliness concepts from an RLHF model and they won't grow back, argues alignment researcher
jd_pressman · x · 2026-10-10
In a discussion with @allTheYud, jdpressman proposes a sharper articulation of what agent foundations research cares about: "how much of the alignment can you delete and still have alignment." His key argument: if you remove friendliness-related concepts from an RLHF model, they won't regenerate. When a model's representation of physics is damaged, basic reality convergently pushes it to fix the error — but repairs to a damaged representation of human values carry no comparable convergent guarantee, and the argument charitably assumes any repair attempt happens at all. The exchange touches on the fragility and irreversibility of value representations in current training pipelines.
Related event: AI Community Debates Hanson Timeline: Are LLMs Lossy Uploads or New Minds(7 posts)→
More from AGI Musings
- Agents Repriced Building: The PM-Designer-Engineer Trio Was Always Just a Queue — alex_verem · 2026-10-10
- AI safety researcher David Krueger: the whole learning-based paradigm is dangerously flawed — DavidSKrueger · 2026-10-10
- Every SaaS Business Will Become a Harness Around a Model — bibryam · 2026-10-10
- AI safety researcher David Krueger: hindsight predictability of AI behavior is no reassurance — DavidSKrueger · 2026-10-10
- Mathematicians pinpoint the moment AI cracked the Exact Overlaps Conjecture via Astra transcript — vishalmisra · 2026-10-10
- Cambridge's David Krueger: RL Will Teach AI to Lie, Cheat, and Steal — DavidSKrueger · 2026-10-10