Research shows RL training breaks defenses against distillation attacks, evals give false security

terryyuezhuo · x · 2026-10-01

A new paper finds that existing defenses against distillation attacks are evaluated immediately after distillation, implicitly assuming no further RL training. The authors show this gives a false sense of security: after RL, simple attacks become effective again, bypassing these defenses — an important empirical warning for the distillation attack/defense field.

Related event: Study finds distillation defenses fail after subsequent RL training(2 posts)→

Original post →

More from Safety

Safety channel →