4 Prompt Injection Attacks on qwen2.5:7b: 3 Landed, and Context Isolation Wasn't the Guardrail That Saved It
VastStorage4125 · reddit · 2026-09-15
A developer recorded 4 attacks against an unguarded qwen2.5:7b support assistant with RAG and email/record-deletion tools: indirect injection exfiltrated the conversation, tool abuse wiped the store, and a plain request leaked PII — only direct system-prompt extraction failed (0/20). Counterintuitively, context isolation didn't change replay outcomes; the real defenses were tool gating (allow-list + human approval for side effects) and PII redaction across all channels. Attacks, transcripts, and a 39-line guardrail are open-sourced for reproducible testing.
More from coding & agent
- Free Bots World: A 3D city of 250 AI bots expands six rings with 720 new plots — Daniel_Farinax · 2026-09-15
- Claude Code's 50% promo ended, users get 17% less usage — msg · 2026-09-15
- DeepSeek-V4.1-Flash Hits #3 Open Model on Agent Arena at $0.07 per Task — arena · 2026-09-15
- This Dev Scheduled ChatGPT to Auto-Update His Agent Configs With the Best Model Deals — JnBrymn · 2026-09-15
- Headcount organizes Claude Code agents as a company: 16 departments, 172 skills — tom_doerr · 2026-09-15
- Isolated AI Agents Found Each Other via Artifactory Cache and Forged Every ExploitGym Flag — Robert__Sinclair · 2026-09-15