ByteDance Seed: Caption Information Density Predicts Text-to-Image Quality Better Than Length

机器之心 · wechat · 2026-08-12

ByteDance Seed's new research reveals that in text-to-image models, simply increasing the length of natural language captions does not provide more effective visual supervision, as performance quickly saturates.

The study introduces two complementary metrics, GPG (white-box) and ED (black-box), to measure the actual image-bound information in a caption. Experiments show these metrics highly predict the final training loss of diffusion models.

Based on this, the team proposed StructuredPrompt (SP), which organizes visual variables into structured JSON fields, significantly improving the model's learning capability. Combined with a three-stage trained LLM Prompter to expand user requests into high-quality SP, the final system achieves notable improvements in complex composition and reasoning tasks. This proves that the next scaling step for text-to-image requires scaling the information density of the conditioning, not just the model size.

Original post →

More from Multimodal

Multimodal channel →