Study finds distillation defenses fail after subsequent RL training
A new study shows that existing defenses against distillation attacks are evaluated only immediately after distillation, assuming no further training. The authors find that subsequent RL training can easily break these defenses, revealing a threat-model mismatch.
2026-09-29 ~ 2026-10-01 · 2 related posts
- Distillation defenses easily break after reinforcement learning, study finds — Shidan Javaheri · 2026-09-29
- Research shows RL training breaks defenses against distillation attacks, evals give false security — terryyuezhuo · 2026-10-01