LoT: Adaptive Token Layout Speeds Up Image Generation Up to 4.6x

GordonWetzstein · x · 2026-10-07

A Stanford team led by George Naka (with Gordon Wetzstein's group) proposes LoT (Layout of Tokens): generation adapts to the layout, with fine tokens allocated where the prompt needs detail, and speedup determined by the layout's token budget. In their example, LoT uses 6,632 tokens instead of 14,336 for full resolution, generating 2.5x faster — and up to 4.6x when detail is more concentrated. Paper and project page are available.

Related event: Stanford's Level-of-Token Diffusion Speeds Up Generation Up to 4.6x(4 posts)→

Original post →

More from Multimodal

Multimodal channel →