Diffusion Models Scale Like LLMs, But Need 10x Data Per Parameter

burny_tech · x · 2026-08-25

The paper 'ABRA: Scaling Diffusion Image Training' reveals that while diffusion image models scale predictably like LLMs, their compute-optimal recipe differs significantly. They require roughly 200 image tokens per parameter (about 10x Chinchilla) and are far more tolerant to overtraining than undertraining. Thus, for fixed compute, training smaller models on more data is the safer bet.

Original post →

More from Models

Models channel →