1.2M ImageNet images + recaptioning rival FLUX at T2I, using 1/1000th the data
serrjoa · x · 2026-09-08
An arXiv paper (2502.21318, Degeorge et al.) challenges the 'bigger is better' paradigm in text-to-image training: with only 1.2M ImageNet images, detailed recaptioning plus CutMix augmentation yields a general T2I model matching FLUX-dev.
- 0.67 GenEval, on par with FLUX-dev; +5 overall over SD3 on GenEval and +12 on DPGBench over SDXL
- Uses just 1/1000th the training images and 3-10x fewer parameters
- Standardized setup needs only 500 H100 hours, and ImageNet is publicly available
The authors argue data design matters more than sheer scale, opening a path to fully reproducible T2I research.
More from Multimodal
- One Prompt Recreates Higgsfield's Animated Short Demo: Blender + ComfyUI + Local MiniMax H3 — Hailuo_AI · 2026-09-08
- Workflow: GPT Astra builds a Blender dungeon, MiniMax H3 renders the walkthrough — Hailuo_AI · 2026-09-08
- Image-to-interactive-3D pipeline: Hyper3D plus Three.js plus AI coding — TheMoonMidas · 2026-09-08
- Users Claim Astra Just Outclassed Every Video Editing Tool — _AustinCalvert_ · 2026-09-08
- Pro animator says H3 anime clips look better than expected, eyes AI shorts — GrungeWerX · 2026-09-08
- Video editor says GPT-6 Astra rebuilt his full Premiere Pro edit in minutes from raw footage — _AustinCalvert_ · 2026-09-08