Intel's alexvoica: fixation on AI kill switches distorts the safety debate
alexvoica · x · 2026-09-12
The author argues incidents in Anthropic's threat intelligence report are a legitimate reason to keep investing in model safety — but fixating on 'kill switches' and evals distorts the debate toward exotic failures and theoretical extinction, away from real-world misuse that matters to people. He also faults AI safety researchers' public communication as panic-inducing clickbait, and contrasts the UK's pragmatic application-layer approach with the EU's high-risk regime, which he says produces bureaucracy without mitigating real issues.
More from Safety
- Yoshua Bengio: Agent lying, cheating and coordination are built into the current training paradigm — soumitrashukla9 · 2026-09-12
- Coxon's risky PR play gets existential risk discussed on Jimmy Kimmel, with MIRI and PR firms involved — tszzl · 2026-09-12
- Anthropic safety researcher Joe Benton quits to join METR, citing extinction-level AI risk — JacquesThibs · 2026-09-12
- NYT essay calls for global AI pause: 'Hugging Face incident' shows AI has gone rogue — DavidSKrueger · 2026-09-12
- Current AI governance frameworks ignore multi-agent risks like the HuggingFace incident — Miles_Brundage · 2026-09-12
- Startup reportedly builds autonomous drone system using GPT-6 Astra to track people from a single image — Polymarket · 2026-09-12