Ex-Omni-2D: Expressive Omni-Modal Dialogue Models with Native Visual Presence
Haoyu Zhang · hf · 2026-08-12
Ex-Omni-2D is an omni-modal dialogue framework capable of generating coordinated text, speech, and video responses. It achieves expressive interaction through a visual thought plan combined with a distilled streaming video generator.
More from Multimodal
- Flux 3 [T2V] Demo: Generating 1990s-Style Monster Footage — CurieuxExplorer · 2026-08-12
- Wan 3.0 Tested: Major Physics Upgrades & 30-Second Clips — Fresh-Resolution182 · 2026-08-12
- Open-Source MiniMax H3 Optimization Suite Cuts VRAM Usage by 25% — Fantastic-Equal-1696 · 2026-08-12
- MiniMax H3 Turbo LoRA Released: 4-Step Generation at 768p — jugernaut126 · 2026-08-12
- Qwen-Image-3.0 Hits OpenArt with Native Text Rendering in 12 Languages — Alibaba_Qwen · 2026-08-12
- Bypass Alibaba Cloud: Calling Wan 3.0 via Aggregator API — Practical_Low29 · 2026-08-12