Counterintuitive: generic synthetic data beats self-generated data for distillation adapters

ostrisai · x · 2026-09-24

Developer ostris found a surprising result training distillation adapters for the Ming Image model: the v1 adapter trained on a generic 50/50 synthetic dataset (done for lack of time) works remarkably well, while v2 trained properly on the model's own generated data performs worse.

His hypothesis: because the model knows its own data, it doesn't break down the distillation hard enough. Self-synthetic adapters have a known limitation of not surviving long training; the generic-data mix apparently doesn't. After 14k steps, the generic-data adapter preserved distillation significantly better. He cautions this is one case and balanced dataset crafting needs more exploration.

Related event: Counterintuitive: generic synthetic data beats model-generated data for distillation(2 posts)→

Original post →

More from Multimodal

Multimodal channel →