Alibaba's Tongyi-MAI: Pixel-Space Diffusion Up to 4.75x Faster

Crazy-Repeat-2006 · reddit · 2026-08-20

Researchers from Alibaba's Tongyi team published a paper on training pixel-space text-to-image diffusion models. Observing that direct large-scale pre-training in pixel space converges slowly, they propose a latent-to-pixel strategy. This approach acquires generative priors in latent space before transitioning to pixel space during post-training. By optimizing key factors like weight initialization and noise schedule, the resulting models match or outperform latent-space counterparts while achieving 3.18x to 4.75x end-to-end inference speedups.

Related event: Tongyi Proposes Latent-to-Pixel Strategy to Speed Up Diffusion Models(2 posts)→

Original post →

More from Research

Research channel →