LLaDA-Image: 6B diffusion model with Muon optimizer hits open-source SOTA image generation

inclusionAI · hf · 2026-09-04

inclusionAI released LLaDA-Image, unifying a 6B diffusion transformer with a frozen vision-language module. Using image-only pre-training and the Muon optimizer, it generates photorealistic images with precise editing, and is distilled into a 2-4 step variant claiming state-of-the-art open-source results, with fully open training recipes.

Related event: Ant's InclusionAI Open-Sources LLaDA-Image Diffusion Models(2 posts)→

Original post →

More from Multimodal

Multimodal channel →