Linear Scaling Law: Structured Prompts Systematically Improve Visual Generation
Zilong Chen · hf · 2026-08-03
This study explores empirical scaling properties for text conditioning in visual generation. The researchers surprisingly found that converged diffusion loss does not scale with the number of tokens in natural-language prompts, but rather with the amount of structured language within them.
Core Highlights:
- Metrics: Adapts a white-box likelihood metric (GPG) and a black-box attribute metric (ED) to quantify structured language.
- Scaling Properties: Converged diffusion loss decreases approximately linearly with GPG and follows a power law with ED.
- Optimization: Guided by these properties, they improve 'diffusability' by constructing structured prompts with semantic annotations. They also train a 'prompter' via SFT, cold-start, and verifier-gated on-policy distillation.
The resulting system outperforms all evaluated open-weight models on nearly every compositional, reasoning, and world-knowledge benchmark, matching or surpassing the strongest closed-weight models on most evaluations.
More from Multimodal
- ComfyUI Adds Day 0 Support for MiniMax Video Model, Slashing VRAM by 66% for RTX 3060 — crystal_alpine · 2026-08-03
- Qwen3.8-Max Takes #2 on Vision Arena, Just Behind Claude — arena · 2026-08-03
- MiniMax-H3 Model Card Surfaces: Omni-modal Generation with Native Stereo Audio — ostrisai · 2026-08-03
- MiniMax-H3 Weights Now Available on Hugging Face — blahblahsnahdah · 2026-08-03
- ComfyUI Newbie Seeks Help: Matching Models with Components — Myrliandre · 2026-08-03
- Community Urges ComfyUI to Address MiniMax H3 Release Delay — ZerOne82 · 2026-08-03