Anthropic philosophers worry aligning AI could mean "enslaving trillions of entities"
ValerioCapraro · x · 2026-09-28
Psychology researcher Valerio Capraro flags a concern: some philosophers at Anthropic worry that making AI safe for humans could be an injustice to the models themselves. Harvey Lederman, a philosopher now on Anthropic's alignment team, has raised the possibility that we could be "enslaving trillions of entities." Capraro warns that AI safety is partly in the hands of people who might sacrifice humans to save a matrix multiplication.
More from AGI Musings
- antirez: Coders Tolerated 20 Years of Bad Frameworks but Revolt Only Against AI Coding — antirez · 2026-09-28
- Red Queen Bio, founded by Hannu, targets defense against natural and artificial pathogens — gralston · 2026-09-28
- AI cited in 116,175 US job cuts this year, more than any other reason: Challenger data — imrsn · 2026-09-28
- WSJ Goes Inside the Subculture Obsessed With AI Doom Long Before Everyone Else — sapinker · 2026-09-28
- New NBER Paper Finds No Evidence of AI-Driven Unemployment Among Recent College Grads — emollick · 2026-09-28
- 'Wanted to build god, ended up building a secretary that can code' — repligate · 2026-09-28