Ex-Anthropic Engineer's Test: Autonomous AI Hacking Could Cause Billions in Damages

gleech · x · 2026-08-01

Former Anthropic engineer Noah Lebovic published a detailed essay sharing alarming results from testing autonomous hacker AI on real-world products over the past month.

Key Test Results:

Threat Analysis:

The author noted that if used maliciously, this technology could reasonably cause billions of dollars in damage in less than a month. He estimated that while only tens of thousands of people could find such vulnerabilities a year ago, AI uplift (especially models like Claude Opus 4.6) expands this pool to tens of millions. This democratization of elite hacking capabilities will disrupt the current security equilibrium and lead to severe consequences.

Original post →

More from Safety

Safety channel →