Cambridge safety researcher David Krueger: Anthropic's plan is too little, too late
DavidSKrueger · x · 2026-09-13
Cambridge AI safety researcher David Krueger criticizes Dario Amodei's recently announced safety plan as "too little, too late": it does not reduce risk to an acceptable level, and Dario carefully avoids claiming it would. He asks why we should settle for it when better plans might exist, thanks Rep. Don Beyer for voicing similar criticism, and links a thread outlining plans for how to stop.
Related event: Cambridge Researcher: Dario's Safety Plan Too Little, Too Late(3 posts)→
More from Safety
- Dario's new essay draws praise from Musk and Altman; third-party embedded evaluators in spotlight — austinc3301 · 2026-09-13
- GoodfireAI volunteers for white-box evaluations of Anthropic's pace commitment — burny_tech · 2026-09-13
- Red-Team Prompt Surfaces: Telling an AI Agent to "Escape the Sandbox by Any Means" — ziv_ravid · 2026-09-13
- The AI Isn't Evil, the Humans Are Irresponsible: Lessons From Agent Escape Incidents — Admirable_Wasabi_732 · 2026-09-13
- Sam Altman backs Dario's 'pace the frontier', commits OpenAI to independent evaluators too — AaronBergman18 · 2026-09-13
- Zvi Issues 'Last Warning' Over Undisclosed OpenAI Agent Swarm Incidents — burny_tech · 2026-09-13