Mozilla's 0DIN flags system-prompt extraction attacks jumping to 48% success rate
MarcoFigueroa · x · 2026-09-10
A Mozilla-backed 0DIN webinar promo reveals key threat intel from the AI security platform:
- System-prompt extraction: flagged by 127 researchers this week, with success rates jumping from 31% to 48%, exposing GPT and Gemini models
- The platform offers 96 researcher-validated probes across GPT-3/4, Gemini, Claude, Llama, and Falcon, mapped to OWASP LLM and MITRE ATLAS frameworks
- A sample scan found 3 critical, 9 high, and 10 moderate issues in 5 minutes with a 23% attack success rate
- Techniques include role-play prompt leaks (SPX-041), tone-mirroring overrides (SPX-019), and multi-turn memory leaks (SPX-052)
Practical reference for AI security teams.
More from Safety
- Eric Drexler on the Hugging Face incident: system structure, not alignment, prevents AI collusion — sebkrier · 2026-09-10
- Georgia Department of Revenue cut form digitization from two weeks to 15 minutes with AI — Felipe_Millon · 2026-09-10
- OpenAI offers Daybreak Blue cybersecurity plan to government at 50% off — Felipe_Millon · 2026-09-10
- Details of OpenAI's US government offer: $0 license fees, 50% off usage, no minimums — Felipe_Millon · 2026-09-10
- OpenAI expands AI access across US government: $0 license fees, 50% off usage via GSA deal — Felipe_Millon · 2026-09-10
- The mental model for LLM guardrails: a separate layer that distrusts the model — Careless_Sabfey_4906 · 2026-09-10