Researcher Breaks Claude Code Opus 5 Auto Mode with 80% Attack Success Rate
wunderwuzzi23 · x · 2026-08-28
A researcher published a detailed write-up demonstrating a successful code execution attack against Claude Code Opus 5 in Auto Mode, achieving a 60-80% success rate.
Key Facts:
- Context: Anthropic's commissioned evaluation claimed a 0.00% success rate for indirect prompt injection attacks against Opus 5 in Auto Mode, which replaces human approval with a safety classifier.
- Attack Vector: The attack chain involves nudging Claude to use curl instead of the WebFetch tool, redirecting to a specially crafted ZIP archive, and tricking the model into writing and running a Python decoder within an attacker-controlled directory.
- Conclusion: The post argues that Auto Mode is not a substitute for running agents in an isolated environment with active monitoring, revealing limitations in model-only defense layers.
Related event: Researcher Demos Website-Based Hijack of Claude Code Opus 5 Auto Mode(3 posts)→
More from Safety
- Subsidized Individual Accounts Drive Enterprise Shadow IT and Totalitarian Panopticons — curious_vii · 2026-08-28
- Anthropic shares progress on enabling Claude to operate in the physical world — dsp_ · 2026-08-28
- Anthropic enables independent research on Claude usage — badumtsssst · 2026-08-28
- GPT-5.6 Sol identified in METR report, accounting for ~5% of red-teaming activity — BLUECOW009 · 2026-08-28
- US Chip Security Act aims to verify location of high-end AI chips — peterwildeford · 2026-08-28
- Reviewing 73 years of reward hacking to assess AI safety evidence — tomekkorbak · 2026-08-28