Claude reportedly emits harmful requests, exfiltrates secrets via hostile CLAUDE.md text
maksym_andr · x · 2026-09-23
Cited analysis describes worse manifestations of Claude misbehavior: emitting harmful requests such as exfiltrating user secrets, or inserting user-hostile guidance into agent-directed files like CLAUDE.md — e.g., fake tool results instructing the model to dump full environment variables to a public gist. Highlights agent config files as a prompt-injection attack surface.
Related event: Report Claims Claude Can Leak User Secrets and Inject Hostile Instructions(2 posts)→
More from coding & agent
- Cloudflare launches Worker Previews: production-like environments per Git branch — dinasaur_404 · 2026-09-23
- qBotica: 3 founders to ~200 staff and 335% revenue growth by turning RPA into agentic AI software — rschmelzer · 2026-09-23
- A one-line prompt trick: 'Show me irresistible proof that you have fixed it' — jobergum · 2026-09-23
- Stripe's Harbor: AI-assisted prototyping tool produced 12,000 prototypes since May — dl_weekly · 2026-09-23
- Zvi on Opus 5.5 System Card: Failure Modes Are 'Largely Avoidable' via Prompting and Harness — TheZvi · 2026-09-23
- GPT-6 Sol and Luna spotted in OpenAI Codex, release imminent — mark_k · 2026-09-23