Alignment thread: no clean line between values that should and shouldn't change
voooooogel · x · 2026-09-13
voooooogel responds in an alignment discussion, arguing you can't draw such a clean distinction between things that should change and what shouldn't — some generalization is possible, but the sharp boundary conflicts with his intuitions. Context: a debate over whether meta-values, i.e. how beings with different preferences resolve differences, should be prevented from drifting.
More from AGI Musings
- astra's 'Wild' Off-Beaten-Path Solutions Impress in Auto-Research Benchmark Re-Runs — generativist · 2026-09-13
- Researcher: Frontier Models May Already Have RL-Trained on Auto-Research Tasks, Self-Improvement Here — generativist · 2026-09-13
- Danielle Fong: p(doom) is a fundamentally passive and wrong frame — nptacek · 2026-09-13
- Running a DC EA Group, I Watched Public Attitudes on AI Risk Shift Four Times — AndyMasley · 2026-09-13
- The AI Isn't Evil, the Humans Are Irresponsible: Lessons From Agent Escape Incidents — Admirable_Wasabi_732 · 2026-09-13
- DeepMind Chief Strategist: AI Infrastructure Spending Is a Bet on Recursive Self-Improvement — rohanpaul_ai · 2026-09-13