Linear Scaling Law: Structured Prompts Systematically Improve Visual Generation
Zilong Chen · hf · 2026-08-03
This study explores empirical scaling properties for text conditioning in visual generation. The researchers surprisingly found that converged diffusion loss does not scale with the number of tokens in natural-language prompts, but rather with the amount of structured language within them.
Core Highlights:
- Metrics: Adapts a white-box likelihood metric (GPG) and a black-box attribute metric (ED) to quantify structured language.
- Scaling Properties: Converged diffusion loss decreases approximately linearly with GPG and follows a power law with ED.
- Optimization: Guided by these properties, they improve 'diffusability' by constructing structured prompts with semantic annotations. They also train a 'prompter' via SFT, cold-start, and verifier-gated on-policy distillation.
The resulting system outperforms all evaluated open-weight models on nearly every compositional, reasoning, and world-knowledge benchmark, matching or surpassing the strongest closed-weight models on most evaluations.
Related event: Structured Prompts Follow Linear Scaling Law in Visual Generation(2 posts)→
More from Multimodal
- Startup claims a neuron-inspired software layer makes AI video 5x faster and 80% cheaper — The Decoder · 2026-09-22
- AI-Generated Short Film Depicts a Cinematic 'Glass' Delivery Rider — NoVeterinarian5438 · 2026-09-22
- One-line prompt turns any animal into a pastel 3D Sherlock Holmes — umesh_ai · 2026-09-22
- ComfyUI Adds Native Loop Nodes: Guide Shows Iterative Video Extension and MLLM-Judged Auto-Editing — nomadoor · 2026-09-22
- Seedance 2.5 Turns AI Influencer Vlogs Into 30-Second Cinematic Stories — aftahi_ai · 2026-09-22
- Seedance 2.5 used to create viral 'iPhone Duo on the street' POV meme video — aftahi_ai · 2026-09-22