3B-parameter model Iris generates every pixel directly, no VAE needed

Deadity · reddit · 2026-10-10

A Reddit user shared speridlabs/iris-3b on Hugging Face: a 3-billion-parameter image generation model that skips the VAE entirely and models every pixel directly. It's an unusual architecture choice versus mainstream latent-diffusion approaches, and an interesting study for anyone exploring non-diffusion generative routes.

Related event: Sperid Labs Open-Sources Iris-3B, a Pixel-Space Text-to-Image Model(8 posts)→

Original post →

More from Multimodal

Multimodal channel →