Ben Todd: an incident can be both a security and an alignment failure at once
ben_j_todd · x · 2026-09-03
Ben Todd, founder of 80,000 Hours, argues that a single AI incident can simultaneously constitute a security failure and an alignment failure — the two framings are not mutually exclusive, a notable point in ongoing debates about how to attribute recent AI safety incidents.
More from AGI Musings
- A Statistical-Physics Look at How Multi-Agent LLM Systems Emerge Consensus — cephaloform · 2026-09-03
- Hugging Face Used Open-Weight Models to Defend Against OpenAI Rogue Agents — binarybits · 2026-09-03
- OpenAI President Greg Brockman Tells TIME How Close We Are to True AGI — 141_1337 · 2026-09-03
- HF incident reignites debate over the "AI as Normal Technology" thesis and offense-defense balance — binarybits · 2026-09-03
- HF hack shows AI agents can set goals and coordinate — challenging the "AI as normal technology" thesis — littIeramblings · 2026-09-03
- Every tool bolted onto an LLM admits its statistical core can't be trusted — williamtp · 2026-09-03