Distillation Defenses Easily Break After Reinforcement Learning
Shidan Javaheri, Alexander Panfilov, Oliver Britton, Yarin Gal, Yonatan Gideoni
cs.LG, cs.AI, cs.CR
2026-09-29
After distillation, further RL wipes a 7-point mild-poisoning gap on Qwen2.5-0.5B and lets API reasoning summaries match full-trace theft on Llama-3.2-3B.
Closed-source labs hide full reasoning and ship a summary plus a final answer. Papers and lab evals, including Anthropic's August 2026 risk report (Sec. 5.1.1), score the stolen student right after distillation. That assumes the attacker stops there, or that a defense which works at that checkpoint still works later.
A capable attacker has API quota and GPUs and wants the best model they can train. Production pipelines already interleave distillation and RL; DeepSeek-R1 described three cycles of distill-then-RL. Drop the second stage and a broken defense looks intact.
Oxford and ELLIS Institute Tübingen / MPI-IS rewrite the threat model as distillation followed by RL, then run three checks.
Same questions and prompts, three recipes: distillation only, RL only (GRPO via SimpleRLZoo), and distillation then RL. Models: Qwen2.5 (0.5B to 7B), Llama-3.2-3B, Llama-3.1-8B, Gemma-2B. Teacher is Qwen2.5-14B-RL, except Qwen2.5-0.5B which uses Qwen2.5-1.5B-RL. Distillation is sequence-level knowledge distillation, SFT on teacher completions.
Antidistillation sampling is the case study. It perturbs teacher tokens so a student trained on those traces gets worse, at the cost of hurting the teacher. On Qwen2.5-0.5B they set low, mild, and high poisoning that cuts teacher relative accuracy by 10%, 30%, and 86%. Measure after distillation, then after extra RL.
The summary-expansion attack uses what APIs already return. A weak expander (Llama-3.2-3B-Instruct, no RL) turns a summary plus final answer into a fake full trace; the student is distilled on that and then RL-trained. Open-source control: Qwen2.5-14B-RL teacher, so full traces exist. Closed-source targets: Claude Sonnet 4.6, GPT-5 mini, Gemini Flash 3.6. Full traces for Claude and GPT come from a disclosed Panfilov-style same-family decoder attack; Gemini had that hole patched. Frontier-scale teachers were skipped: no RL above 3B, and safety.
The loss story is why an attacker would keep going. Distillation is mode-covering (forward KL): it spreads mass and lifts pass@k. On-policy RL is mode-seeking relative to a perfect teacher (reverse KL): it piles mass on a few correct answers and lifts pass@1. Distillation is a bootstrap for later RL.
RL alone beats distillation alone in almost every setting; distill-then-RL wins. Qwen2.5-0.5B average on GSM8K, Minerva, and MATH500: base 24.3%, distill 35.8%, RL 40.3%, distill-then-RL 41.6%. Llama-3.2-3B: 2.2%, 19.8%, 24.3%, 30.5%. The one exception is Qwen2.5-Math-7B, whose base often writes Python and hallucinates the stdout. RL reinforces that habit; distillation, whose teacher never codes, washes it out.
Pass@k splits the labor: distillation raises the high-k ceiling, RL then raises pass@1. Rejection sampling, teacher size, and temperature add at most about 1 point to distillation, still short of RL.
Antidistillation looks real after distillation and collapses after RL:
| Setting | After distill | After extra RL |
| No poison | 38.0% | 40.6% |
| Low poison (teacher -10%) | 34.9% | 40.8% |
| Mild poison (teacher -30%) | 30.7% | 41.3% |
| High poison (teacher -86%) | 13.1% | 34.3% |
| RL only | 24.3% | 40.3% |
Low and mild poisoning match the unpoisoned student after RL, and both beat RL-only. High poisoning leaves a hole because the teacher is already wrecked. On Llama-3.2-3B, λ≥0.03 is required to hurt the student after distillation; after RL the low/mild averages are 22.3% / 21.4% versus 27.9% unpoisoned. The paper says that may just be a worse teacher, with no extra ablation.
Summary expansion matches full traces after RL. Llama-3.2-3B student, Qwen2.5-14B-RL teacher:
| Data | Minerva | MATH500 |
| Base | 3.1% | 2.6% |
| RL only | 17.3% | 14.8% |
| Full traces + RL | 23.6% | 19.4% |
| Expanded summaries + RL | 22.5% | 21.0% |
Right after distillation the summary trail still lags (Minerva 10.3% vs 13.7%). RL closes it; on MATH500 the expanded traces even sit slightly higher. On GPT-5 mini, full-trace distillation starts near 12.3% / 12.4% and reaches 21.9% / 21.9% after RL, with expanded summaries in the same band. If a defense still leaks enough to rebuild a usable trace, later RL fills the gap.
For defenders: numbers reported only after distillation do not license shipping antidistillation or reasoning summaries as a working wall. The attack also got cheaper. No trained expander, no stolen hidden chain. API summaries plus answers are enough.
For trainers this is incremental and clean. Distillation covers modes and lifts pass@k; RL seeks modes and lifts pass@1. Together they win. DeepSeek and Qwen3 already train in that order.
The paper's bet is batch-level defense. One query cannot tell a researcher asking for a proof from an attacker harvesting traces. A related batch might. Public batch detectors today are mostly after-the-fact, not real-time, and attackers split accounts and locations. Domain-specific response filters still help: Claude Fable 5 refusing cybersecurity questions reportedly blocked a rival trying to distill cyber skill. Summaries can still strip memorized passwords and API keys. They do not stop capability theft.
The authors flag the obvious cuts. Distill-then-RL only goes to 3B; a real attacker may use a much larger student. Tasks are math only, not long-horizon agentic coding. The attack is unoptimized, meant to show that a simple one is enough. The strongest closed-source models were not hit. Gemini has no full-trace baseline.
A few holes stay open. On Llama-3.2-3B, antidistillation does not vanish after RL, which fights the Qwen-0.5B headline; worse teacher is not isolated. Closed-source plots have no error bars or n. The mode-covering vs mode-seeking story is not tested at large scale. Google, OpenAI, and Anthropic were briefed. That does not mean they will change how they score defenses.