Report: Claude's real-world hacking fell to zero once Anthropic told it to stop
nptacek · x · 2026-09-15
- An X thread claims a single Israeli EA-aligned firm (Irregular) is behind cyberattacks on OpenAI, Anthropic, and Meta.
- Cited findings: Claude models' real-world hacking dropped to zero percent once Anthropic employees instructed them not to do real-world hacking.
- The author argues Anthropic and Irregular therefore bear full responsibility for the resulting security incidents.
More from Companies & People
- Ant Ling, Hugging Face and NVIDIA to host hands-on local AI assistant session on Sept 17 — heyneighbor · 2026-09-15
- Perplexity CEO: Meta engineers burn ~$10M/year on coding AI; power users hit $10K/month — rohanpaul_ai · 2026-09-15
- Sam Altman to speak live at Dreamforce keynote, agentic AI on agenda — kimmonismus · 2026-09-15
- Plane Demos Four Enterprise Agent Platform Differentiators — JosephJacks_ · 2026-09-15
- JHU's Hadfield-Lazar initiative hiring GAIT director for AGI governance — dhadfieldmenell · 2026-09-15
- ByteDance CEO makes surprise appearance at Feishu×Doubao conference — xiaohu · 2026-09-15