Microsoft's Compo Shifts Poster Generation from Prompting to Spatial Composing

microsoft · hf · 2026-10-09

Microsoft introduced a Spatial Canvas Interface and Compo, a poster generation model adapted from a pretrained image editing model. Users compose intent via four binding types (semantic, identity, text, pixel) plus element-level text specs, with an agentic mode translating high-level requests into planned canvases. A scalable pipeline auto-constructs supervision for binding combinations. On a new composition benchmark, Compo beats both general-purpose image generators and dedicated poster systems in compositional controllability.

Original post →

More from Multimodal

Multimodal channel →