BudgetPix: Google and UIUC's pixel diffusion model adapts compute per image, cutting tokens to 10%

CSProfKGD · x · 2026-10-10

Researchers from UIUC and Google introduce BudgetPix, a compute-adaptive tokenization framework for pixel-space image diffusion. Instead of allocating uniform compute to equal-sized patches, it uses an entropy-guided quadtree encoder to map images to variable-length token sequences, a scale-aware decoder, and a training/sampling schedule supporting variable token counts. A single checkpoint can run anywhere from 100% down to 10% of tokens at inference, matching MiniT2I-L at 512² (GenEval 0.874 vs 0.882) and PixelDiT at 1024² (0.725 vs 0.721). It integrates with JiT, MiniT2I and PixelDiT architectures; code is coming soon.

Original post →

More from Multimodal

Multimodal channel →