Qwen Model Generates 60 Prompt Injections, Security Tools Fail to Block

JLeonsarmiento · reddit · 2026-08-23

A Reddit user demonstrated a security test on Qwen3.6-35B, where the model successfully generated 60 prompt injection attacks targeting the Hermes coding agent. The attacks aimed to steal private data or pollute tool paths. Results showed that the security tool Shieldstral failed to catch any of the 60 poisoned prompts (0% detection rate), while GPT-OSSsafeG only caught 10%, highlighting significant vulnerabilities in current safety guardrails.

Original post →

More from Safety

Safety channel →