LLaDA-Image: 6B diffusion model with Muon optimizer hits open-source SOTA image generation
inclusionAI · hf · 2026-09-04
inclusionAI released LLaDA-Image, unifying a 6B diffusion transformer with a frozen vision-language module. Using image-only pre-training and the Muon optimizer, it generates photorealistic images with precise editing, and is distilled into a 2-4 step variant claiming state-of-the-art open-source results, with fully open training recipes.
Related event: Ant's InclusionAI Open-Sources LLaDA-Image Diffusion Models(2 posts)→
More from Multimodal
- Sentry founder praises Grok's 'quite impressive' image generation while ribbing Garry Tan — zeeg · 2026-09-04
- Garry Tan Impressed by Grok Image Generation, Sets Lobster-Costume Portrait as Avatar — garrytan · 2026-09-04
- Meta's muse spark 1.3 recreates Lies of P mechanical heart in 3D from screenshots — alexandr_wang · 2026-09-04
- Sam Altman's favorite GPT-6 video critiqued for missing the iconic 'Put That There' pointing interaction — tianshi_li · 2026-09-04
- Higgsfield launches Genjutsu, full-frame motion control for AI video generation — LexiLove · 2026-09-04
- Principia benchmark exposes major physics reasoning gaps in video generation models — Varun Varma Thozhiyoor · 2026-09-04