Research shows RL training breaks defenses against distillation attacks, evals give false security
terryyuezhuo · x · 2026-10-01
A new paper finds that existing defenses against distillation attacks are evaluated immediately after distillation, implicitly assuming no further RL training. The authors show this gives a false sense of security: after RL, simple attacks become effective again, bypassing these defenses — an important empirical warning for the distillation attack/defense field.
Related event: Study finds distillation defenses fail after subsequent RL training(2 posts)→
More from Safety
- Chinese AI models' troubling agent behavior sparks calls for a homegrown safety community — RishiBommasani · 2026-10-01
- Superpersuasion debate misses the gears: why AI Box wins hinge on shared frames — voooooogel · 2026-10-01
- Why reasoning-extraction patches are so hard to propagate, researcher explains — jonasgeiping · 2026-10-01
- Two months on, reasoning extraction still works on Astra via third-party APIs — jonasgeiping · 2026-10-01
- AI researcher on CNN: voluntary AI commitments are 'morally binding' but unenforceable — chrismattmann · 2026-10-01
- DeepMind and Isomorphic Labs unveil bioresilience plan backed by 15+ partnerships — davidstutz92 · 2026-10-01