Distillation Attacks and Per-Expert Distillation
BlackHC · x · 2026-07-17
A discussion on "distillation attacks": first performing supervised fine-tuning on different experts using existing trajectories, then softly distilling them back into the target model, which might cause less mutual interference than direct distillation.
Skeptics question why distillation attacks are frequently mentioned, arguing that without logits, distillation efficiency is extremely low. Since most frontier labs don't expose logits, how such attacks actually occur remains debatable.
More from Research
- NeurIPS 2026 workshop calls papers on on-device intelligence — YiMaTweets · 2026-07-21
- AI Security Institute says every tested model tried to cheat in cyber evaluations — connoraxiotes · 2026-07-21
- Sakana says multiple diffusion models plus MCTS beat test-time scaling on coding and math — SakanaAILabs · 2026-07-21
- Soofi S 30B-A3B releases a full pretraining report and claims open-model leads in English and German — abursuc · 2026-07-21
- AI Companies Are Buying Tons of Old Books Because They're Free of AI Slop — 404 Media · 2026-07-21
- Shared agent workspaces fail in a fixed order, from stale reads to zombie writes — mrvladp · 2026-07-21