Counterintuitive distillation finding: generic synthetic data beats model-generated data for adapters

ostrisai · x · 2026-09-24

ostrisai found a surprising result while distillation-training adapters: for the Ming Image V1 adapter, he had no time to generate a proper dataset, so he trained on a generic 50/50 synthetic dataset — and it worked remarkably well. The V2 adapter, trained properly on images generated by the model itself as he normally does, performed worse. He says H3 is the first model he'll retrain on generic data.

Related event: Counterintuitive: generic synthetic data beats model-generated data for distillation(2 posts)→

Original post →

More from Multimodal

Multimodal channel →