New 13,000-word essay: AI safety should bet on control and governance over alignment
sayashk · x · 2026-09-15
Sayash Kapoor and Arvind Narayanan (authors of AI as Normal Technology) published a 13,000-word essay analyzing recent loss-of-control incidents at AI companies and how to "pace the frontier." Key arguments:
- Bridge the two camps: AI safety researchers frame incidents as an alignment crisis; cybersecurity practitioners see failed basic hygiene. A middle-ground approach is the way forward.
- AI control isn't solved: OpenAI's protections were inadequate, and known control methods would have prevented the Hugging Face incident — but as agents improve, control requires sustained investment.
- Specific risks over generic doom: these incidents are primarily a security story; cyberoffense is the urgent autonomous-agent risk, and biorisk and military AI deserve targeted defenses too.
- Marginal investments in control beat alignment: companies neglect available control techniques, and common-sense policy can push investment there.
- Organizational governance is the key lever: technology alone can't fix governance if reckless teams can simply opt out of using controls.
More from AGI Musings
- OpenAI Capabilities Researcher Dan Selsam Publishes Personal Statement on AI Risk — connoraxiotes · 2026-09-15
- 'People Who Say We Need to Nuke SF Are More Hypocritical Than OpenAI' — wordgrammer · 2026-09-15
- e/acc's Beff Jezos Argues Capability Diffusion Is Safest, Slams Anthropic's Closed Approach — beffjezos · 2026-09-15
- The User-Assistant Format Is an Illusion: Why Persona-Based AI Alignment Likely Won't Work — mayfer · 2026-09-15
- Clarifying the AI safety split: safety testing time vs long internal deployment — JacquesThibs · 2026-09-15
- a16z partner puts P(abundance) at 99.99% — nptacek · 2026-09-15