LLaDA-Image: 6B fully-diffusion DiT trained on 90% image-only data, no caption bottleneck
jiqizhixin · x · 2026-09-16
Inclusion AI released LLaDA-Image, a from-scratch 6B fully-diffusion DiT where both the backbone and generation are diffusion models. Key insight: it learns visual priors from images alone—over 90% of 220M cumulative training samples use image-only supervision without captions, reserving image-text pairs for later language alignment, sidestepping the scarce/expensive caption bottleneck. One set of weights supports multiple tasks.
More from Models
- Mozilla Report: China-US AI Model Capability Gap Narrows to 4.4 Months — External_Mood4719 · 2026-09-16
- Dev recalls fighting for structured output in GPT-3 days as dynamic inputs now parse in <200ms — blixt · 2026-09-16
- llama.cpp Merges hc Ops for Qwen4exp, Time to Re-benchmark Qwen Flash — jacek2023 · 2026-09-16
- Apodex 1.1 demos: raw clinical tables to Kaplan-Meier curves, XRD indexing, protein docking — aakashgupta · 2026-09-16
- Apodex 1.1 launches: moves from deep research to real task execution — testingcatalog · 2026-09-16
- Indie dev open-sources Qwen-2.5-1B-RLCD with 5x faster on-device JSON inference — suchenzang · 2026-09-16