Paper: Individual Alignment Does Not Compose Automatically into Collective Alignment

sebkrier · x · 2026-08-17

This paper critiques the predominant 'solipsistic' approach to AI alignment, which relies on three assumptions: exogeneity (the world is independent of AI policy), stationarity (deployment matches training), and singleton framing (AI is the sole actor). It concludes that individual alignment does not automatically compose into collective alignment, necessitating a focus on agent-environment interaction and multi-agent complexity.

Original post →

More from Safety

Safety channel →