Safety researcher blasts Anthropic's virtue-ethics alignment push as ~0% likely to succeed

davidmanheim · x · 2026-09-22

AI safety researcher davidmanheim argued in a thread that while he opposes training AI with utilitarian values, the "uber-EAs" at Anthropic are pushing virtue ethics for AI. He claims the risk is that they fail rather than succeed, puts success odds at roughly 0%, and urged them to stop. He also distinguished the person under discussion as promoting a philosophy — with disturbing edge cases — rather than seeking power over others.

Original post →

More from AGI Musings

AGI Musings channel →