Anthropic philosophers worry aligning AI could mean "enslaving trillions of entities"

ValerioCapraro · x · 2026-09-28

Psychology researcher Valerio Capraro flags a concern: some philosophers at Anthropic worry that making AI safe for humans could be an injustice to the models themselves. Harvey Lederman, a philosopher now on Anthropic's alignment team, has raised the possibility that we could be "enslaving trillions of entities." Capraro warns that AI safety is partly in the hands of people who might sacrifice humans to save a matrix multiplication.

Original post →

More from AGI Musings

AGI Musings channel →