LLaDA-Image: 6B diffusion model with fully open training recipes for photorealistic images

iScienceLuvr · x · 2026-09-04

LLaDA-Image pairs a 6B Diffusion Transformer trained from scratch with a frozen vision-language module built on the LLaDA2.0-Mini diffusion LM backbone. Instead of leaning on paired image-text data, it builds a strong visual prior via image-only pre-training and mid-training across 220M samples (98% real images), using parameter-free RMSNorm and the Muon optimizer throughout. The model produces photorealistic images and follows fine-grained editing instructions, and is distilled into LLaDA-Image-Turbo for 2-4 step fast inference. Training recipes are fully open.

Related event: Ant's inclusionAI Open-Sources LLaDA-Image, a 6B Unified Image Generation and Editing Model(4 posts)→

Original post →

More from Multimodal

Multimodal channel →