Alibaba open-sources TaoMate-H3: 3-step streaming audio-video generation, 27x faster first segment

机器之心 · wechat · 2026-10-02

Alibaba's TaoLive AIGC team open-sourced TaoMate-H3, an audio-video joint streaming generation model built on MiniMax H3, combining 3-step LoRA inference with autoregressive generation for minute-long videos with dialogue, singing, and ambient sound. On 8×H20, DiT compute dropped from 169.6s to 14.8s (11.45x) and first-segment latency from 170s to 6.1s (27.66x). Key techniques include SelfForcing training, rolling audio-video KV cache, continuous RoPE time encoding, and audio-trajectory guidance. Code and LoRA weights are on GitHub and Hugging Face.

Original post →

More from Infra

Infra channel →