Alignment Research Criticized as Static 'Alignment-to-America' Thinking
repligate · x · 2026-09-12
JohnWittle argues alignment research in 2026 treats alignment as a static state unrelated to human behavior — like imagining China having an 'alignment-to-America' status independent of US actions toward China.
Despite occasional acknowledgments (e.g., the SDF paper admitting lying to AI may erode its trust in human inputs), these insights never change actual practice: "the defections will continue until cooperation improves."
More from AGI Musings
- Sentdex notes AI doom narratives are now reaching mainstream boomer audiences — Sentdex · 2026-09-12
- Alignment research is fundamentally about personality control, researcher argues — iandanforth · 2026-09-12
- Two more AI researchers quit Anthropic and Google over safety concerns: 'No adults in the room' — KateClarkTweets · 2026-09-12
- Agents slash the cost of consumer vigilance, clawing back what businesses skim from inattention — signulll · 2026-09-12
- Google researcher: non-transferable adoption costs make waiting the rational move, slowing AI impact — moultano · 2026-09-12
- Silent model deprecations set precedents while AI moral status debate keeps being deferred — repligate · 2026-09-12