LLaDA-Image: 6B diffusion model with fully open training recipes for photorealistic images
iScienceLuvr · x · 2026-09-04
LLaDA-Image pairs a 6B Diffusion Transformer trained from scratch with a frozen vision-language module built on the LLaDA2.0-Mini diffusion LM backbone. Instead of leaning on paired image-text data, it builds a strong visual prior via image-only pre-training and mid-training across 220M samples (98% real images), using parameter-free RMSNorm and the Muon optimizer throughout. The model produces photorealistic images and follows fine-grained editing instructions, and is distilled into LLaDA-Image-Turbo for 2-4 step fast inference. Training recipes are fully open.
More from Multimodal
- Gemini API launches Agentic Video: up to 88% fewer tokens and better long-video reasoning — davidstutz92 · 2026-09-04
- Early user: Seedance 2.5 delivers effortlessly cinematic shots, a favorite gen AI model — heypearlai · 2026-09-04
- Minimax-h3-Turbo ships FL2V Turbo 4-step v1.2 at 768p — Any_Fee5299 · 2026-09-04
- Local Dune skit on a 4090: character sheet plus voice reference workflow — r0ni · 2026-09-04
- Using Qwen3-VL-4B to auto-expand prompts for Minimax H3 video generation — CountFloyd_ · 2026-09-04
- Z-Image Base prompting experiment: natural-language scene blocks beat tag lists — Maleficent-Bowl-4841 · 2026-09-04