SEMA trains open-source models for multi-turn adversarial attacks without human scripts
burkov · x · 2026-09-30
Security researcher Andriy Burkov introduces SEMA, a framework addressing multi-turn jailbreak threats.
- Malicious users can gradually obscure harmful intent across dialogue turns; existing attacks suffer from high exploration complexity and intent drift
- SEMA trains open-source LLMs to generate effective multi-turn adversarial attacks
- Requires no human-crafted scripts, predefined strategy templates, external jailbreak corpora, or closed-source APIs
More from Safety
- Researcher Points to Transluce's Approach as a Model for Verifying Lab Usage Claims — agstrait · 2026-09-30
- Anthropic Report Says GLM-5.3 Safeguards Bypassed 64%-100% of the Time, Drawing Open-Source Backlash — robleclerc · 2026-09-30
- Researchers Question Reliability of Frontier Labs' Self-Reported Usage Statistics — agstrait · 2026-09-30
- AI news roundup: AISI supply-chain attack report, AMD's $8.2B World Labs deal, OpenAI DevDay — Justgototheeffinmoon · 2026-09-30
- Altman at DevDay: more resources for safety, IPO now 'ill-advised', says Dots best in market — shiringhaffary · 2026-09-30
- Ethan Mollick: Open-weights models will soon pose the same security threats as closed ones, without guardrails — emollick · 2026-09-30