ORCA Paper: Single Training-Only Loss Fixes Text-to-Image Composition Failures

mo_lotfollahi · x · 2026-10-09

ORCA (accepted at NeurIPS 2026) tackles compositional failures in text-to-image models: ask for "a red cube next to a blue ball" and colours bind to the wrong object, spatial relations flip, and counts drift — each object looks perfect, but composition breaks.

Part of the cause is the text encoder: CLIP reads prompts almost like a bag of words, knowing which concepts appear but not how they relate. SD3 and FLUX added a T5 encoder to preserve structure, yet the failures persist. The authors' view: the information is there, but nothing in training teaches the model to tie it to the image.

ORCA adds the missing signal with a single training-only loss at zero inference cost. On DiT-L/2 it beats vanilla and REPA models trained twice as long, with 3x gains on colour binding and spatial relations.

Original post →

More from Multimodal

Multimodal channel →