Intel's alexvoica: fixation on AI kill switches distorts the safety debate

alexvoica · x · 2026-09-12

The author argues incidents in Anthropic's threat intelligence report are a legitimate reason to keep investing in model safety — but fixating on 'kill switches' and evals distorts the debate toward exotic failures and theoretical extinction, away from real-world misuse that matters to people. He also faults AI safety researchers' public communication as panic-inducing clickbait, and contrasts the UK's pragmatic application-layer approach with the EU's high-risk regime, which he says produces bureaucracy without mitigating real issues.

Original post →

More from Safety

Safety channel →