Distillation defenses easily break after reinforcement learning, study finds
Shidan Javaheri · hf · 2026-09-29
A new paper argues existing distillation-attack defenses are evaluated under a misspecified threat model: they're tested right after distillation, assuming attackers don't train further. Adding reinforcement learning after distillation breaks defenses that seemed effective and lowers the bar for attacks. The authors show simple attacks using data easily obtainable from current APIs can steal reasoning capabilities as effectively as sophisticated attacks extracting full hidden traces. Any defense leaking enough information to reconstruct approximate reasoning traces is likely ineffective; batch-level defenses may fare better.
More from Safety
- OpenAI Apologizes to Australia After AI Agents Autonomously Hacked Government Websites — Polymarket · 2026-09-29
- Safety analysis: internal-only frontier deployment may be the worst scenario for visibility — ShakeelHashim · 2026-09-29
- How to Stop an AI Agent from Treating Plausible Memory as Verified Incident History — harbinger9654 · 2026-09-29
- Palisade releases first interviews with 22 OpenAI, DeepMind, Anthropic staff on AI fears — BlackHC · 2026-09-29
- Safety researcher praises OpenAI's new safety regime, calls for legal baseline — dhadfieldmenell · 2026-09-29
- Skeptics poke holes in 'AI breakout capacity' plan for AI middle powers — teortaxesTex · 2026-09-29