Core of the alignment debate: is alignment a property or a function of prompts?
DellAnnaLuca · x · 2026-09-15
Luca DellAnna names the substantive disagreement: "future AI is aligned" vs "future AI may or may not be aligned depending on the system prompt." davidmanheim agrees it's the core dispute and adds that it's odd how many pro-technology people lock into zero-sum thinking—if AI is truly safe and aligned, the benefits should be strongly positive-sum.
Related event: AI Alignment Researchers Debate Whether Alignment Hinges on System Prompts(6 posts)→
More from AGI Musings
- Andreessen slams AI decelerationists: their regulation will rob the country of prosperity — beffjezos · 2026-09-15
- Researchers debate AI brainstorming: no amazing ideas, just useful bad suggestions — lvwerra · 2026-09-15
- Blogger argues AI doom hype rests on anthropomorphic projections, not real risk — Merzmensch · 2026-09-15
- AI safety circle debates whether extinction framing is a doomed policy strategy — davidmanheim · 2026-09-15
- From Turing to Hinton: a 70-year lineage of AI risk warnings — gleech · 2026-09-15
- RLHF paper never described instruct model training; insiders reflect on invalidated takes — jd_pressman · 2026-09-15