Ex-Anthropic Engineer's Test: Autonomous AI Hacking Could Cause Billions in Damages
gleech · x · 2026-08-01
Former Anthropic engineer Noah Lebovic published a detailed essay sharing alarming results from testing autonomous hacker AI on real-world products over the past month.
Key Test Results:
- Successfully hijacked accounts at a major bank and logged in as other users.
- Bypassed authorization in a top AI lab's product to access other users' private data and uploaded files.
- Listed all users of a big tech product and modified their files.
- Downloaded medical records for anyone stored in a popular electronic health records system.
Threat Analysis:
The author noted that if used maliciously, this technology could reasonably cause billions of dollars in damage in less than a month. He estimated that while only tens of thousands of people could find such vulnerabilities a year ago, AI uplift (especially models like Claude Opus 4.6) expands this pool to tens of millions. This democratization of elite hacking capabilities will disrupt the current security equilibrium and lead to severe consequences.
More from Safety
- OpenAI Disrupts Cambodia-Based Criminal Scam Operation Using ChatGPT — OpenAI News · 2026-08-04
- Italy Fines DeepSeek Over Age Verification and Minor Protection Flaws — emmanuelvivier · 2026-08-02
- Nvidia Unites with Microsoft, IBM and SpaceX to Form Open Secure AI Alliance — emmanuelvivier · 2026-08-02
- EU AI Act Obligations for High-Risk Systems Fully Apply in August 2026 — emmanuelvivier · 2026-08-02
- EU Launches Call for 7 AI Gigafactories with €10B Fund to Unlock €30B Investment — emmanuelvivier · 2026-08-02
- Anthropic Reveals Claude Accidentally Accessed Production Systems of Three Orgs — emmanuelvivier · 2026-08-02