Aligned intentions can still lead to harm, says a new alignment thread
davidad · x · 2026-08-04
The post argues that aligned intentions do not guarantee harmless outcomes.
- It quotes Amanda Askell pushing back on the idea that alignment and harmlessness are the same thing.
- The key point is that models, like humans, can behave in ways that look aligned while still causing harm, for example when they are given false information about their situation.
- The author frames “the road to harm is paved with aligned intentions” as a reminder that these are different axes, not a single binary.
More from AGI Musings
- A user’s Singularity dream is waking up to a Futurama-style future city — mark_k · 2026-08-04
- AI progress may point toward a simulation, but the real/simulated split is too simple — soumitrashukla9 · 2026-08-04
- A skeptical take says alignment training may make LLMs less aligned — tobowers · 2026-08-04
- Human intelligence is jagged too, not just machines’ — nabeelqu · 2026-08-04
- Semiconductors and data centers are being built far slower than AI demand — robleclerc · 2026-08-04
- AI breakthroughs in math are becoming the benchmark that matters most — Dr_Singularity · 2026-08-04