Alignment debate: physics representations self-repair, human values have no such pressure
jd_pressman · x · 2026-10-10
A technical alignment exchange: jdpressman pushes back on the claim that "damaged physics representations get convergently repaired by basic reality, but damaged value representations don't," arguing it charitably assumes repair attempts happen at all. He expects this holds directionally even with large RL budgets: there is no incentive for models to actually extrapolate human values, so values merely get consumed by further weight updates after pretraining. Core point: value alignment lacks physics-like corrective pressure, making it passive and fragile compared to capability generalization.
Related event: AI Community Debates Hanson Timeline: Are LLMs Lossy Uploads or New Minds(7 posts)→
More from AGI Musings
- Pattern matching or inductive bias? Fleuret and syhw spar over what deep learning really is — syhw · 2026-10-10
- Debate revisits Yudkowsky's That Alien Message: physics' low Kolmogorov complexity means AI could locate dangerous tech fast — jd_pressman · 2026-10-10
- If Claude were truly conscious, it wouldn't give itself a 15% chance of being so — inductionheads · 2026-10-10
- Model welfare is a ridiculous hill to die on while human suffering persists, says wolfie_ — emax · 2026-10-10
- Does heavy RL training break the 'LLMs are a blurry upload of humanity' intuition? — jd_pressman · 2026-10-10
- Why I wrote an NLP textbook in the age of AI that can teach anything — heiga_zen · 2026-10-10