TILT improves compositional text-to-image generation with a model-intrinsic reward
Debottam Dutta · hf · 2026-07-29
TILT proposes a training-free way to improve compositional text-to-image generation
- The paper, TILT: Improving Compositional Generation in Diffusion Models with a Model-Intrinsic Reward, targets a common failure mode in text-to-image models: handling complex prompts with multiple concepts.
- It introduces a training-free test-time method that changes the sampling trajectory using a reward derived from the base model itself, so it does not need external supervision or a separate reward model.
- The authors frame compositional failures as overlap between joint and single-concept distributions, then derive a KL-constrained objective with a closed-form tilted target distribution and guidance steps for diffusion sampling.
- On T2ICompBench, the method improves compositional alignment while preserving image quality, and a hybrid guidance strategy performs best among the explored variants.
More from Multimodal
- Developer turns ChatGPT Real-time Voice into a desktop 3D persona — BLUECOW009 · 2026-07-29
- Qwen Audio 3.0 Realtime Plus costs $4.42 per input-audio hour in testing — ArtificialAnlys · 2026-07-29
- Qwen Audio 3.0 Realtime Plus tops Speech-to-Speech benchmark at 84.1% — ArtificialAnlys · 2026-07-29
- AI black-comedy short film turns a zombie apocalypse into a music video — Aggravating-Chest997 · 2026-07-29
- Claude 5 Opus generates a textureless dirt-road car demo entirely on its own — ChrisGPT · 2026-07-29
- Claude 5 Opus turns a no-texture dirt-road car demo into fully generated game graphics — ChrisGPT · 2026-07-29