Training a 210M text-to-image DiT on one GPU in 3.5 days: three measured findings
IvanMikhnenkov · reddit · 2026-09-11
The author trained a 210M cross-attention DiT text-to-image model from scratch on a single RTX PRO 6000 (3.5 days, 4.2M images at 256²) to understand the full recipe. Three findings rarely stated plainly:
- Learned null attention slots become the sink: 2 learned K/V slots receive 90% of cross-attention mass in middle blocks at mid-noise; EOS drops to 4%; register vectors grow to 4–13× the norm of image tokens.
- Flow-matching loss is a health signal, not quality: loss moved only 0.805→0.754 while held-out FID went 33.7→27.0, FD-DINOv2 570→218, object accuracy 65%→90%.
- Training-time timestep shift beats doubling steps: 20 steps with shift 2.8 → FID 27.0; 50 steps → 26.6; no shift → 27.3 and FD-DINOv2 228.
Setup: 896×16 blocks, 2D RoPE, QK-norm, adaLN-single, rectified flow with logit-normal timesteps, frozen flan-t5-base, batch 256 for 400k steps, torch.compile giving 2.4× speedup. Code, weights and demo are open-sourced; next step is Flow-GRPO on this base.
More from Multimodal
- Ethan Mollick prompts Google's Astra to build a game about the end of all things — eldonredwards · 2026-09-11
- Cristóbal Valenzuela unveils Solaris, an "interface world model" — c_valenzuelab · 2026-09-11
- World's first AI idol Yuri holds a concert with a stunning opening — xiaohu · 2026-09-11
- AI short film "The World Worth Living In" enters global AI film festivals — Xanavibes · 2026-09-11
- Same Prompt, Same Seed: Anima vs SDXL Image Quality Compared — nano_chad99 · 2026-09-11
- Synthesia Adds Custom Avatars Built From a Text Prompt — synthesiaIO · 2026-09-11