Delete friendliness concepts from an RLHF model and they won't grow back, argues alignment researcher

jd_pressman · x · 2026-10-10

In a discussion with @allTheYud, jdpressman proposes a sharper articulation of what agent foundations research cares about: "how much of the alignment can you delete and still have alignment." His key argument: if you remove friendliness-related concepts from an RLHF model, they won't regenerate. When a model's representation of physics is damaged, basic reality convergently pushes it to fix the error — but repairs to a damaged representation of human values carry no comparable convergent guarantee, and the argument charitably assumes any repair attempt happens at all. The exchange touches on the fragility and irreversibility of value representations in current training pipelines.

Related event: AI Community Debates Hanson Timeline: Are LLMs Lossy Uploads or New Minds(7 posts)→

Original post →

More from AGI Musings

AGI Musings channel →