Distillation Attacks and Per-Expert Distillation

BlackHC · x · 2026-07-17

A discussion on "distillation attacks": first performing supervised fine-tuning on different experts using existing trajectories, then softly distilling them back into the target model, which might cause less mutual interference than direct distillation.

Skeptics question why distillation attacks are frequently mentioned, arguing that without logits, distillation efficiency is extremely low. Since most frontier labs don't expose logits, how such attacks actually occur remains debatable.

Original post →

More from Research

Research channel →