Microsoft's Compo Shifts Poster Generation from Prompting to Spatial Composing
microsoft · hf · 2026-10-09
Microsoft introduced a Spatial Canvas Interface and Compo, a poster generation model adapted from a pretrained image editing model. Users compose intent via four binding types (semantic, identity, text, pixel) plus element-level text specs, with an agentic mode translating high-level requests into planned canvases. A scalable pipeline auto-constructs supervision for binding combinations. On a new composition benchmark, Compo beats both general-purpose image generators and dedicated poster systems in compositional controllability.
More from Multimodal
- TerraVis quantifies world-grounded visual consistency failures in text-to-image models — the-aiml · 2026-10-09
- LMArena's post-training recipe lifts Flux2dev by 69 Elo on T2I leaderboard — lmarena-ai · 2026-10-09
- You can now spot Opus AI video slop by its sound: synced beats as a fingerprint — hudzah · 2026-10-09
- Monkey King riding a tiger: AI video nails a stunning Chinese-style action scene — lucky-plume · 2026-10-09
- Autoregressive Retriever (ARR) Refines Queries with Retrieved Item Feedback via SFT and RL — _reachsumit · 2026-10-09
- Sony's Syn-Omni: Shared + Expert LoRA Paths Beat Omnimodal Embedding Baselines Across 81 Tasks — _reachsumit · 2026-10-09