Alignment debate: models are context-dependent distributions, so alignment is impossible
gerardsans · x · 2026-09-18
Replying to @stevesi, @gerardsans argues that alignment rests on a false assumption — that there is a fixed something to align. A model is a mathematical distribution whose behavior changes with context rather than a fixed trait, so alignment under current technology is impossible. He adds that it "can't be buggy because it's probabilistic — it behaves exactly how it was designed."
More from AGI Musings
- Runway CEO: the frontier is moving from generating content to generating worlds — c_valenzuelab · 2026-09-18
- Charity Majors publicly rejects podcast invite written with ChatGPT: 'reads as an insult' — mipsytipsy · 2026-09-18
- Could a specialized superintelligence solve the alignment problem? A Reddit proposal — fullcongoblast · 2026-09-18
- Intelligence supercycle relied on capital-formation innovation, not just tech, VC argues — inductionheads · 2026-09-18
- Recursive training collapses LLMs by gen 9; 10% human data halts the damage — alex_verem · 2026-09-18
- Ex-computational linguist: mathematicians' reaction to LLMs puts his old field to shame — voooooogel · 2026-09-18