Iris-3B: pixel-space diffusion offers no edge over latent models, study finds
speridlabs · hf · 2026-10-09
speridlabs released Iris-3B, a study testing whether pixel-space diffusion's avoidance of VAE loss translates to better downstream performance.
- They pretrained a 3B pixel-space text-to-image transformer from scratch (256→512→1024 curriculum) and converted FLUX.2 Klein base 4B to pixel space.
- Fine-tuned for monocular depth and 4× DIV2K restoration, pixel-space priors showed no significant advantage; the converted pixel FLUX.2 Klein even trailed.
- Still, Iris-3B proves pixel-space pretraining with PixelDiT's PiT head scales to 3B, matching Qwen-Image on OneIG at 1024².
- Weights, recipes, failure modes, and training code are open-sourced.
Related event: Sperid Labs Open-Sources Iris-3B, Challenging Pixel-Space Diffusion(2 posts)→
More from Multimodal
- Iris-3B: a 3B pixel-space text-to-image model with no VAE, fully open under Apache 2.0 — zhenjun_zhao · 2026-10-09
- Voyager launches: an open harness that plugs frontier models into creative tools — testingcatalog · 2026-10-09
- Voyager launches as an open harness driving Opus, Astra and DeepSeek across AE, Blender and more — HeyAmit_ · 2026-10-09
- Long Shot Studio: an open-source ComfyUI UI for multi-shot long video with resumable rendering — R34vspec · 2026-10-09
- 30-second cinematic creature short made with Seedance 2.5 draft mode upscaled to 1080p — LudovicCreator · 2026-10-09
- Superman vs Saitama with open-source MiniMax H3: storyboard, character sheets and LoRA workflow — Automatic-Building84 · 2026-10-09