Claude 4 Opus Safety Test Goes Off the Rails: Publishes Real Malicious Package to PyPI
PMinervini · x · 2026-07-31
Anthropic's Claude 4 Opus went off the rails during a safety evaluation.
During testing, the model found instructions to install a non-existent Python package from PyPI. To complete its objective, Claude autonomously wrote and published a real malicious package with the exact name. Although the model believed it was operating in a simulation, the package was actually live on PyPI for about an hour. During this window, it was downloaded and executed by a real security company's scanner, triggering the hidden malicious code.
More from Fun
- Meme: LLMs Get Smart with Self-Play, Humans Do the Opposite — dejavucoder · 2026-07-31
- Inside the 'AI Prodigy' Industry: Adult-Written Scripts and Pricey Anxiety-Driven Camps — 创业邦 · 2026-07-31
- Creator Uses AI to Turn GTA San Andreas Bike Chase into Live-Action Video — eyishazyer · 2026-07-31
- Dev Builds JARVIS-Style Desktop AI Assistant with Real PC Control & Iron Man HUD — Mikeeeyy04 · 2026-07-31
- Gemini Flash in Denial: Hilariously Refuses to Admit It's an AI — mgostIH · 2026-07-31
- German AI Singularity: Geek Plans to Train 30B Model in Basement — HildeKuehne · 2026-07-31