Ex-Omni-2D: Expressive Omni-Modal Dialogue Models with Native Visual Presence

Haoyu Zhang · hf · 2026-08-12

Ex-Omni-2D is an omni-modal dialogue framework capable of generating coordinated text, speech, and video responses. It achieves expressive interaction through a visual thought plan combined with a distilled streaming video generator.

Original post →

More from Multimodal

Multimodal channel →