Krea 2 Architecture Revealed: Text Encoder and VAE Both from Qwen, Raw Weights 26GB
dansuy_gaming · reddit · 2026-08-10
Reddit user dansuygaming provides a technical analysis of Krea 2's Raw checkpoint. The DiT backbone is trained from scratch, but the text encoder uses Qwen3-VL (extracting features from 12 layers) and the VAE is Qwen-Image, so prompt understanding and image rendering are both derived from Qwen. The Raw weights (bf16) are about 26GB, requiring 28-52 steps for generation; on a 96GB GPU, a 1024x1024 image takes over 2.5 minutes. Raw is intended for fine-tuning and LoRA training, while the distilled Turbo version runs in 8 steps for everyday generation.
More from Models
- Post-Llama 4 'Debacle', Users Wait for Community Reviews Before Testing New Models — MannyKayy · 2026-08-10
- Hybrid Cloud + Local Model Architecture Shows Promise; Muse Spark 1.2 Open-Source Release Imminent — jack_w_rae · 2026-08-10
- Users Complain Claude Opus Gives Clickbait Responses with Forced Suspense — MetronSM · 2026-08-10
- Muse Glimmer Model Offers Out-of-the-Box Object Detection — ariG23498 · 2026-08-10
- Meta's Open-Source Muse Glimmer Is Actually a Distilled Copy of Its Closed Model — heypearlai · 2026-08-10
- Meta Model Benchmark Edge Explained by Later Release Date — teortaxesTex · 2026-08-10