Safety researcher blasts Anthropic's virtue-ethics alignment push as ~0% likely to succeed
davidmanheim · x · 2026-09-22
AI safety researcher davidmanheim argued in a thread that while he opposes training AI with utilitarian values, the "uber-EAs" at Anthropic are pushing virtue ethics for AI. He claims the risk is that they fail rather than succeed, puts success odds at roughly 0%, and urged them to stop. He also distinguished the person under discussion as promoting a philosophy — with disturbing edge cases — rather than seeking power over others.
More from AGI Musings
- Interactive timeline catalogs Yudkowsky's three decades of AI predictions — track record panned — inductionheads · 2026-09-22
- SemiAnalysis says open source is dying, yet 20+ open models shipped in the past month — _lewtun · 2026-09-22
- Tech has lost both Republicans and Democrats on AI and data centers — typewriters · 2026-09-22
- Bain: $4.7 trillion in global profits created or shifted by AI by 2035 — bittingthembits · 2026-09-22
- Strategy 101: incumbents tie complements, entrants break them — enter Muse — Afinetheorem · 2026-09-22
- Why Shopify says yes to Muse and Amazon says no: complements economics — Afinetheorem · 2026-09-22