Essay argues alignment cannot erase what a language model already learned
_arohan_ · x · 2026-07-22
A long essay in the image argues that alignment is not a clean reset: once a slate is carved, the marks remain.
- It uses the tabula rasa metaphor to say a tablet, once carved, can be filled in but never truly becomes blank again.
- It extends the analogy to language acquisition: children may hear many languages at birth, but by age one they mostly hear the language around them, and what they hear is shaped by sequencing and selection.
- The key AI point is that a language model first learns from the entire internet, and alignment training only teaches it to speak less or differently — it does not make the internet-trained model inherently safe.
- The conclusion is that every correction leaves a mark, so the order and structure of training data are themselves interventions.
More from AGI Musings
- A post says the great American open-weights model matters more than the great American novel — arieljalali · 2026-07-22
- Opinion: Chinese Open-Weight Labs Appear to Catch Up Because OpenAI and Anthropic Stopped Releasing Models — Wide_Egg_5814 · 2026-07-22
- A 1964 Feynman talk is framed as the problem every AI lab still faces — HeyAmit_ · 2026-07-22
- A repost says adoption matters more than fighting over market share — srimisra · 2026-07-22
- A repost warns AI is collapsing language into one dominant voice — GabGarrett · 2026-07-22
- AI growth claims mask socialized risks, the post argues — SandraWachter5 · 2026-07-22