Alignment debate: physics representations self-repair, human values have no such pressure

jd_pressman · x · 2026-10-10

A technical alignment exchange: jdpressman pushes back on the claim that "damaged physics representations get convergently repaired by basic reality, but damaged value representations don't," arguing it charitably assumes repair attempts happen at all. He expects this holds directionally even with large RL budgets: there is no incentive for models to actually extrapolate human values, so values merely get consumed by further weight updates after pretraining. Core point: value alignment lacks physics-like corrective pressure, making it passive and fragile compared to capability generalization.

Related event: AI Community Debates Hanson Timeline: Are LLMs Lossy Uploads or New Minds(7 posts)→

Original post →

More from AGI Musings

AGI Musings channel →