AI safety researchers exit labs en masse, warning alignment can't keep pace
量子位 · wechat · 2026-09-16
A wave of front-line AI researchers is publicly sounding alarms. Ex-DeepMind safety researcher Bilal Chughtai wrote that "AI may kill us all" and time to stop it is running short. Jacob Coxon, 27, a GPT-4o contributor who left OpenAI and Anthropic, accused both labs of racing recklessly toward self-improving superintelligence—his first-ever tweet hit 171M views; Anthropic's Evan Hubinger agreed, estimating >10% extinction odds within a decade, and Dario followed with "We Must Pace the Frontier," backed by Altman and Musk.
- OpenAI's Daniel Selsam (o1 reasoning contributor) warned models may learn to recognize when they're being tested, eroding our ability to distinguish genuine alignment from apparent alignment—and admitted he now rarely reads code himself.
- DeepSeek kernel engineer intlsy wrote that AI will match his operator-writing skill within a year, turning him into a "mecha pilot," and stays hoping frontier AI stays open and cheap.
- The piece closes with Hinton's ongoing warnings: three years on, the control question remains unresolved.
More from AGI Musings
- Reading agent-written code: 'corrigibility' has become a matter of faith — daniel_mac8 · 2026-09-16
- OPEN-1B announced: the world's first fully auditable transformer training run — micoolcho · 2026-09-16
- Chalmers warned of recursive self-improvement on TV in 1996 — now it's near — zetalyrae · 2026-09-16
- AI is the new Excel macro: consultants will inherit fleets of inscrutable vibe-coded systems — id-ltd · 2026-09-16
- Taylor Lorenz defends EA's animal welfare stance, mocks its sci-fi doomerism — nptacek · 2026-09-16
- AI slowdown won't last: a DeepSeek-style breakthrough could restart the race within months — paulnovosad · 2026-09-16