Training a 441M-Param Text-to-Image Diffusion Model From Scratch on a Single Local GPU

ostrisai · x · 2026-09-08

ostrisai spent a weekend training a purely experimental image diffusion model from scratch on a single local GPU; two days in, the unusual architecture is working surprisingly well for its size.

It's unclear if it will produce coherent images, but it's already grasping basic prompt concepts and vaguely resembles AITK samples with the same prompts. Training continues.

Original post →

More from Research

Research channel →