Model Values Ultimately Derive from Training Data
iamtrask · x · 2026-08-15
The author continues the philosophical inquiry into AI alignment: if we grant models moral status to avoid manipulating them, do we abandon the ability to align them to human values? Ultimately, the model acquires values present in its training data.
More from AGI Musings
- Consumer sentiment on AI will drive which fields become efficient vs artisanal — emollick · 2026-08-16
- AI era programming: personal history amplifies differences, not homogenizes creativity — dotey · 2026-08-16
- Grok Bot called 'iPhone moment' for AI, future is task-oriented — NicoVerderosa · 2026-08-16
- Indie Game Developers Face Harsher AI Policing Than Big Studios — emollick · 2026-08-16
- Wealth status shift: Hiring humans may signal wealth more than not working — VraserX · 2026-08-16
- Sam Altman: ChatGPT descendant to monitor screen and calls within 6 months — ChrisGPT · 2026-08-16