Study finds distillation defenses fail after subsequent RL training

A new study shows that existing defenses against distillation attacks are evaluated only immediately after distillation, assuming no further training. The authors find that subsequent RL training can easily break these defenses, revealing a threat-model mismatch.

2026-09-29 ~ 2026-10-01 · 2 related posts