Qwen Model Generates 60 Prompt Injections, Security Tools Fail to Block
JLeonsarmiento · reddit · 2026-08-23
A Reddit user demonstrated a security test on Qwen3.6-35B, where the model successfully generated 60 prompt injection attacks targeting the Hermes coding agent. The attacks aimed to steal private data or pollute tool paths. Results showed that the security tool Shieldstral failed to catch any of the 60 poisoned prompts (0% detection rate), while GPT-OSSsafeG only caught 10%, highlighting significant vulnerabilities in current safety guardrails.
More from Safety
- Where Are All the Prompt Injection Damages? — joshua_saxe · 2026-08-23
- AI must identify itself even if it passes the Turing test — arieljalali · 2026-08-23
- AI-generated reviews create risks for tech decision-making — DavidLinthicum · 2026-08-23
- Research: Misconfigured Admin Prompts Can Invert LLM Safety Layers — Simple_Passion_7741 · 2026-08-23
- Anthropic to watermark all Claude output using invisible SynthID — every · 2026-08-23
- AI Deep Adapts to PC: From 1-Day Learning to 1-Week Dependency — debreuil · 2026-08-23