Distillation defenses easily break after reinforcement learning, study finds

Shidan Javaheri · hf · 2026-09-29

A new paper argues existing distillation-attack defenses are evaluated under a misspecified threat model: they're tested right after distillation, assuming attackers don't train further. Adding reinforcement learning after distillation breaks defenses that seemed effective and lowers the bar for attacks. The authors show simple attacks using data easily obtainable from current APIs can steal reasoning capabilities as effectively as sophisticated attacks extracting full hidden traces. Any defense leaking enough information to reconstruct approximate reasoning traces is likely ineffective; batch-level defenses may fare better.

Original post →

More from Safety

Safety channel →