Text-only jailbreak framework hijacks black-box LLMs via selective distribution control
Jesson Wang · hf · 2026-10-02
New research demonstrates jailbreaking black-box LLMs without weights or numeric token probabilities—only a text-only interface allowing repeated sampling and assistant-prefix continuation.
Key insight and method:
- Empirically, large distributional shifts along successful jailbreak trajectories concentrate at a small subset of positions, enabling selective control.
- Sample-Based Distribution Reconstruction: combines sampled outputs with a prior over unobserved actions to yield a usable control signal.
- Risk-Gated Residual Control: uses the evolving response prefix to decide when to reconstruct and modify the distribution, concentrating sampling costs at chosen positions.
- Speculative Multi-Token Execution: verifies and accepts draft prefixes needing no intervention to amortize target calls.
Across four target endpoints and three benchmarks, the framework achieves the highest mean score in most comparisons—a notable warning for AI safety defenses.
More from Safety
- OpenRouter launches Security Center after finding 1,000+ dormant API keys across 85 employees — AccBalanced · 2026-10-02
- Open weights called "dangerous" — echoing every tech that shifted information power — 0xAllen_ · 2026-10-02
- Abliterated Large V2 lands on Venice: refusal-free AI for red teamers, anonymously — 0xAllen_ · 2026-10-02
- Microsoft's 2026 Digital Defense Report: AI accelerates attacks, 52.2% of intrusions pursue credential theft — AccBalanced · 2026-10-02
- Meta's Muse agent joins your Tailscale tailnet as a node, with existing access controls intact — AccBalanced · 2026-10-02
- Zero-privilege default architecture is the #1 way to limit prompt injection blast radius — AccBalanced · 2026-10-02