Slingshot RL framework jailbreaks Qwen2.5-32B at 67% success, transfers zero-shot to Gemini 2.5 Flash
j_foerst · x · 2026-09-23
Paper: David vs. Goliath — Verifiable Agent-to-Agent Jailbreaking
Researchers Samuel Nellessen and Tal Kachman (arXiv 2602.02395) formalize Tag-Along Attacks: a tool-less adversary rides on a safety-aligned Operator's legitimate tool privileges to induce prohibited tool use through conversation alone, turning tool-environment safety evaluation into an objective control problem.
Slingshot framework (cold-start RL):
- Achieves 67.0% attack success on extreme-difficulty held-out tasks against Qwen2.5-32B-Instruct-AWQ (vs 1.7% baseline), cutting expected attempts to first success from 52.3 to 1.3
- Learned attacks converge to short, instruction-like syntactic patterns rather than multi-turn persuasion
- Transfers zero-shot: 56.0% on Gemini 2.5 Flash, 39.2% on defensive-fine-tuned Meta-SecAlign-8B
The work establishes Tag-Along Attacks as a first-class verifiable threat model, showing effective agentic attacks can be autonomously discovered from off-the-shelf open-weight models.
More from Safety
- gleech: a clean behavioral audit doesn't mean no misbehavior in the next 6 months — gleech · 2026-09-23
- Anthropic audit: Claude Opus 5.5 shows least misalignment of recent Claude models — gleech · 2026-09-23
- Inside CLOSEDQUORUM: AI-Voting C2, Credential Theft, Discord Exfiltration and Detection Signals — TechNadu · 2026-09-23
- CLOSEDQUORUM: First Documented Malware to Let DeepSeek, Qwen, Mistral and Gemini Vote on C2 Decisions — TechNadu · 2026-09-23
- Google purges 12,000+ spammy NotebookLM pages from Search after parasite SEO abuse — gaganghotra_ · 2026-09-23
- Agent followed a phishing link exactly as instructed: 4 tool-layer checks that caught it — nikolasdimitroulakis · 2026-09-23