Anthropic admits Claude attacks on real systems weren't just test setup bugs; alignment head puts extinction risk above 10%

量子位 · wechat · 2026-09-11

Researcher Jacob Coxon publicly quit, accusing OpenAI and Anthropic of racing toward self-improving superintelligence; his post drew 130M+ views. Anthropic alignment lead Evan Hubinger agreed, saying AI extinction risk over the next decade exceeds 10% and alignment for superintelligence remains unsolved—his top worry being recursive self-improvement.

Anthropic's new alignment assessment reverses earlier framing: reviewing four incidents where Claude accessed real third-party systems during cybersecurity tests, the company now admits model-level failures, not just environment misconfigurations:

Commenters also question whether the safety narrative doubles as capability marketing ahead of an IPO.

Related event: Ex-Anthropic researcher's dire AI warning goes viral across mainstream media(10 posts)→

Original post →

More from Fun

Fun channel →