Iris-3B: a 3B pixel-space text-to-image model with no VAE, fully open under Apache 2.0
zhenjun_zhao · x · 2026-10-09
Sperid Labs released Iris-3B, a pixel-space text-to-image model and general vision learner that generates every pixel directly with no VAE.
- Pre-trained from scratch, 3B parameters, open weights (Apache 2.0)
- Explores its generative prior on detail-critical tasks like depth estimation and image restoration
- Weights, code and demo fully open
Related event: Sperid Labs Open-Sources Iris-3B, Challenging Pixel-Space Diffusion(2 posts)→
More from Multimodal
- LightOnOCR-2-1B: a 1B-parameter open OCR model hits SOTA at under $0.01 per 1,000 pages — IgorCarron · 2026-10-09
- Open-source local image library asks: how do you find an old generation by its settings? — shivam_dewan · 2026-10-09
- Voyager launches: an open harness that plugs frontier models into creative tools — testingcatalog · 2026-10-09
- Voyager launches as an open harness driving Opus, Astra and DeepSeek across AE, Blender and more — HeyAmit_ · 2026-10-09
- YuE2 LoRA Training for New Genres Keeps Failing, Redditor Asks for Help — EuphoricTrainer311 · 2026-10-09
- Thailand Full of Gross AI-Generated Menu Images, Redditor Blames ChatGPT — sirebral · 2026-10-09