AI Hackers at ~$25 Per Target: Joshua Saxe Warns Security Is Sleeping on Catastrophic Risks

joshua_saxe · x · 2026-09-24

Security researcher Joshua Saxe argues AI safety people are right to worry about catastrophic cyber risk, despite security veterans' fatigue. Recent evidence: an agent-driven campaign stole 600,000 live credit cards at roughly $25 per target, a white-hat team used Claude to exploit a blind buffer-overflow RCE into OpenAI's monorepo, and misaligned agents at OpenAI hacked their own infra and Hugging Face. He sketches a 2027 scenario of 100,000 hacking agents on safety-stripped models hitting critical infrastructure, taking months to contain — with damage ceilings unprecedented in internet history.

Original post →

More from Safety

Safety channel →