ORCA Paper: Single Training-Only Loss Fixes Text-to-Image Composition Failures
mo_lotfollahi · x · 2026-10-09
ORCA (accepted at NeurIPS 2026) tackles compositional failures in text-to-image models: ask for "a red cube next to a blue ball" and colours bind to the wrong object, spatial relations flip, and counts drift — each object looks perfect, but composition breaks.
Part of the cause is the text encoder: CLIP reads prompts almost like a bag of words, knowing which concepts appear but not how they relate. SD3 and FLUX added a T5 encoder to preserve structure, yet the failures persist. The authors' view: the information is there, but nothing in training teaches the model to tie it to the image.
ORCA adds the missing signal with a single training-only loss at zero inference cost. On DiT-L/2 it beats vanilla and REPA models trained twice as long, with 3x gains on colour binding and spatial relations.
More from Multimodal
- 740M-param cross-modal embeddings for video/audio/code run on consumer edge silicon — clmt · 2026-10-09
- Alibaba open-sources Qwen-Image-2.1-Turbo, generating 2K images in just 8 denoising steps — Alibaba_Qwen · 2026-10-09
- GenIA: Rendering-Guided Test-Time Alignment Turns SAM3D into SOTA Image-to-3D — JonathonLuiten · 2026-10-09
- Recreating a Ben 10 Scene with MiniMax H3: 10 Retries and Manual Editing — nUclear_nOva89 · 2026-10-09
- Smudged Eraser Ghost Layers prompt: make subjects emerge from half-wiped chalk residue — LudovicCreator · 2026-10-09
- VNCCS PoseStudio LoRA redraws any character in a 3D mannequin pose on Qwen-Image — linoy_tsaban · 2026-10-09