Cisco Research: DeepSeek R1 100% Jailbreak Success Rate, Critical Security Flaws
joshrogin · x · 2026-08-20
Cisco's Robust Intelligence, in collaboration with the University of Pennsylvania, assessed DeepSeek R1's security. Using 50 random prompts from the HarmBench dataset, they conducted automated jailbreak attacks covering six harmful categories. The results showed a 100% attack success rate, failing to block any harmful prompt, starkly contrasting with other leading models. The research suggests DeepSeek's cost-efficient training methods may have compromised safety.
More from Safety
- Sam Altman on the AI dilemma: trade-offs between loss of control and power centralization — r0ck3t23 · 2026-08-24
- Debate erupts over lethal military robots vs. failing civilian units — teortaxesTex · 2026-08-24
- Only 1 of 20 Potential Presidential Candidates Answered AI Pause Query — DavidSKrueger · 2026-08-24
- Chinese Transforming Robot Dog Sparks US Trade Policy Criticism — TinfoilTricorn · 2026-08-24
- Turkey blocks at least 12 Grok posts on national security grounds — Unusual_Variation293 · 2026-08-24
- Nature Comment: Provenance, not interpretability, grounds trust in autonomous science — gabepgomes · 2026-08-24