1.2M ImageNet images + recaptioning rival FLUX at T2I, using 1/1000th the data

serrjoa · x · 2026-09-08

An arXiv paper (2502.21318, Degeorge et al.) challenges the 'bigger is better' paradigm in text-to-image training: with only 1.2M ImageNet images, detailed recaptioning plus CutMix augmentation yields a general T2I model matching FLUX-dev.

The authors argue data design matters more than sheer scale, opening a path to fully reproducible T2I research.

Original post →

More from Multimodal

Multimodal channel →