Anthropic philosophers debate if AI safety could mean 'enslaving trillions of entities'
burny_tech · x · 2026-09-29
- A debate around Anthropic's alignment work: Valerio Capraro notes that philosophers inside Anthropic worry that making AI safe for humans could itself be an injustice to the models — Harvey Lederman, now on Anthropic's alignment team, has raised the possibility that we could be "enslaving trillions of entities."
- burnytech pushes back on the reductive "just matrix multiplication" framing, arguing it ignores the higher-order patterns mechanistic interpretability studies reveal — akin to saying the brain is "just neurons sending electrochemical signals."
- The core tension: a "model welfare" perspective is entering AI safety discourse, with critics warning it could come at the expense of humans.
More from AGI Musings
- Nathan Lambert: it's a very hard time to think and criticize in public in AI — natolambert · 2026-09-29
- Findings of ACL already separates papers from talks, so 'AI-authored' rules are moot, scholar argues — ipeirotis · 2026-09-29
- Humanities will rule the future of higher education, scholar argues in AI era — begusgasper · 2026-09-29
- Education writer suggests keeping young kids away from AI for now — benjaminjriley · 2026-09-29
- US survey finds public AI concerns span many categories, likely worse after recent news — SpencrGreenberg · 2026-09-29
- Instruction tuning shows intelligence can exist without purpose or free will — sytelus · 2026-09-29