BudgetPix: Google and UIUC's pixel diffusion model adapts compute per image, cutting tokens to 10%
CSProfKGD · x · 2026-10-10
Researchers from UIUC and Google introduce BudgetPix, a compute-adaptive tokenization framework for pixel-space image diffusion. Instead of allocating uniform compute to equal-sized patches, it uses an entropy-guided quadtree encoder to map images to variable-length token sequences, a scale-aware decoder, and a training/sampling schedule supporting variable token counts. A single checkpoint can run anywhere from 100% down to 10% of tokens at inference, matching MiniT2I-L at 512² (GenEval 0.874 vs 0.882) and PixelDiT at 1024² (0.725 vs 0.721). It integrates with JiT, MiniT2I and PixelDiT architectures; code is coming soon.
More from Multimodal
- One prompt turned OpenAI's 722 math papers into a 92-second animated explainer via Opus — FinanceYF5 · 2026-10-10
- Dev recreates fictional ChatGPT UIs for Windows 98, XP, Vista and 7 — Midnight_Sun_BR · 2026-10-10
- MiniMax H3 removes people from videos inside ComfyUI — RobbaW · 2026-10-10
- Free Female Face Prompt Builder adds natural-language-to-SD prompt conversion — dobelmont · 2026-10-10
- Model tuning could stop AI video from fighting 2D animation style, says Andrew Carr — andrew_n_carr · 2026-10-10
- Beginner cheat sheet for ComfyUI/MiniMax video generation, '83% less slop' — Nimblecloud13 · 2026-10-10