LLaDA-Image: 6B fully-diffusion DiT trained on 90% image-only data, no caption bottleneck

jiqizhixin · x · 2026-09-16

Inclusion AI released LLaDA-Image, a from-scratch 6B fully-diffusion DiT where both the backbone and generation are diffusion models. Key insight: it learns visual priors from images alone—over 90% of 220M cumulative training samples use image-only supervision without captions, reserving image-text pairs for later language alignment, sidestepping the scarce/expensive caption bottleneck. One set of weights supports multiple tasks.

Original post →

More from Models

Models channel →