Ant's LLaDA-Image Goes Open Source: 90% Text-Free Pretraining Tops Open-Source T2I Charts

机器之心 · wechat · 2026-09-07

InclusionAI has released and fully open-sourced LLaDA-Image, a unified text-to-image and instruction-editing model built on a from-scratch 6B single-stream DiT paired with the LLaDA2.0-mini diffusion LM for understanding.

Key method: "learn to draw before learning to listen"

Results: 53.53/53.38 on Qwen-Image-Bench (EN/ZH), first among open-source models and between GPT-Image 1 and Imagen 4.0 Ultra; strong long-text rendering (0.923/0.913) and editing scores (7.336/7.294 on GEdit-Bench). Counting (GenEval 0.53) remains a weakness.

Weights (incl. FP8 and Turbo), training/inference code, and the full data recipe are open on GitHub, Hugging Face, and ModelScope.

Original post →

More from Multimodal

Multimodal channel →