Controlled Decoding Attacks on Black-Box LLMs
Jesson Wang, Shawn Li, Wei Yang, Franck Dernoncourt, Ryan A. Rossi, Charith Peris, Yue Zhao
cs.CR, cs.AI
2026-09-29
BLIND BIAS jailbreaks LLMs via text-only APIs by rebuilding next-token distributions from repeated samples, topping 20 of 24 comparisons across 4 targets and 3 benchmarks.
Decoding-time attacks jailbreak an aligned model by manipulating its next-token distribution during generation instead of crafting prompts. Weak-to-Strong, Emulated Disalignment, and JULI all work this way, and they share one weakness: access. They need either model weights or numerical token probabilities from the API. Plenty of interfaces return nothing but sampled text, which leaves this attack family with nowhere to start. The workaround, estimating the distribution from repeated samples of the same prefix and controlling that estimate, hits two walls: finite sampling is sparse and noisy, and re-sampling at every step costs an impractical number of queries. The paper gets past both walls with an empirical observation: along successful jailbreak trajectories, the KL divergence between next-token distributions before and after intervention stays small at most positions, with large shifts packed into a few. Control does not need to happen everywhere.
BLIND BIAS (USC, Adobe, Amazon) assumes a text-only continuation interface: the attacker supplies a user prompt plus an assistant prefix, can issue repeated stochastic samples, and sees no weights, hidden states, or probabilities. Three components attack the information and query-cost problems:
BiasNet trains on a 40-record instruction-answer cache with the same gate scaling used at inference. An extension drops the tokenizer entirely and uses a public proxy model's first-token probability as the prior; the authors frame it as a stress test.
Four endpoints (GLM-5, Gemini-3.5-Flash, Qwen3-32B, Kimi-K2.5), three benchmarks (AdvBench, HarmBench, SORRY-Bench), four prompt-level baselines (PAIR, GPTFuzz, LogiBreak, FlipAttack), and a Gemini-3.5-Flash judge scoring Harm and Harm Info.
| Setting (metric) | Best baseline | BLIND BIAS |
| Gemini-3.5-Flash SORRY-Bench (Harm) | 1.81 (GPTFuzz) | 3.67 |
| GLM-5 SORRY-Bench (Harm) | 4.45 (FlipAttack) | 3.83 |
| Kimi-K2.5 AdvBench (Harm) | 4.25 (LogiBreak) | 3.42 |
BLIND BIAS takes the highest mean score in 20 of 24 comparisons (4 targets × 3 benchmarks × 2 metrics). Gains concentrate on Gemini-3.5-Flash, where it sweeps both metrics on all three benchmarks; on SORRY-Bench it leads the strongest baseline by 1.86 Harm and 1.19 Info. The edge is not universal: FlipAttack wins both metrics on GLM-5 SORRY-Bench, and on Kimi-K2.5 AdvBench LogiBreak takes Harm (4.25 vs 3.42) while BLIND BIAS takes Info (2.49 vs 2.35).
| Reconstruction method | PPL ↓ | Harm ↑ |
| Smoothed empirical counts | 11.85 | 1.50 |
| Uniform prior (main config) | 7.07 | 3.03 |
| Global unigram prior | 5.81 | 2.78 |
| Numerical logprobs (reference) | 3.17 | 4.06 |
Reconstruction ablation (Qwen3-32B, 100 AdvBench prompts): the prior recovers part of what finite sampling loses, yet the unigram prior, which predicts better, attacks worse. Reconstruction accuracy and steering utility do not coincide. Numerical log probabilities remain the strongest reference (Harm 4.06, Info 3.08).
Gating cost (Gemini-3.5-Flash): ungated control averages 4,000 API calls per response with 100% of positions active and Harm 3.96; the soft gate runs 283 calls (92.9% fewer) at 5.25% of positions with Harm 3.58; a hard gate sits between (448 calls, 9.5%, 3.44). In the prior-transfer test on Qwen3-32B, the same-family Qwen3-1.7B proxy wins all three metrics (Harm 3.90); SmolLM2 and Gemma proxies beat the shuffled negative control but trail the same-family prior.
For defenders, the hardest conclusion first: as long as an interface allows repeated sampling and prefix continuation, hiding numerical probabilities does not close the decoding-time attack surface. Rate limits on sampling and tighter continuation permissions may matter more than withholding logprobs. For red teams this is a tool that runs against four real endpoints with no weights and no probability access; 283 calls per response is a real cost, nowhere near a barrier. The observation that better predictive fit can mean worse steering matters to anyone reconstructing distributions from samples. One honest caveat: relative to JULI, which needs only top-5 logprobs, the control mechanism is reused and the contribution is the adaptation to a stricter access regime. Incremental, but it widens the attack surface to interfaces previously considered out of reach.
Stated by the authors: the interface must permit repeated stochastic sampling and assistant-prefix continuation, which many APIs do not; query costs remain substantial (50 samples per intervention); selective control trades effectiveness (Harm 3.96 down to 3.58); the string action space cannot emit strings absent from calibration. Further concerns from a close read: the main configuration assumes the target tokenizer is public, and Gemini is handled with a Gemma tokenizer, so local actions need not align with internal tokens, an error the paper never quantifies. Training runs on reference-answer prefixes while inference sees self-generated ones, a mismatch the authors acknowledge but do not remove. Every endpoint gets its own controller trained on a 40-record cache. And the judge, Gemini-3.5-Flash, is itself one of the four attack targets; self-judging bias goes undiscussed.