TaoMate: Alibaba & NJU Framework Solves Error Accumulation in Long Digital Human Videos
量子位 · wechat · 2026-08-02
Alibaba's Taotian Group and Nanjing University proposed TaoMate, an anchor-guided persistent memory framework for joint audio-video generation, aiming to solve recursive error accumulation and computational overhead in long digital human video generation.
Core Technical Approach
- Memory Separation & Anchor Constraints: Separates history into a bounded active context and a fixed-capacity persistent memory. It uses immutable visual anchors to constrain dynamic state updates, effectively suppressing color drift and appearance degradation.
- Decoupled Retrieval & Modulation: Uses residual side attention to retrieve audio-visual history and introduces Reference-AwareFiLM for channel-level appearance modulation, ensuring long-term consistency.
- Causal Distillation & Parallel Inference: Trains a student model with mixed rollout lengths to handle accumulated errors. It utilizes stage-parallel inference to achieve a 35.0 FPS full pipeline output on 3 GPUs.
Experimental Results
On a long-duration evaluation benchmark, TaoMate significantly improved LipSync and LongConsistency. Its DiT throughput reached 1.56x that of the next best method. The team also built a real-time interactive prototype supporting streaming generation, verifying its deployment potential for live digital humans.
More from Multimodal
- Describe your dream world to an AI dragon, which generates the planet for you — repligate · 2026-08-24
- Using kintsugi texture to fix cracks in edited 3D meshes — repligate · 2026-08-24
- Generating Hannibal Character Videos with FL2VA Model — Nimblecloud13 · 2026-08-24
- MiniMax H3 Revives Medieval Short Stories: Complete Workflow Shared — zanatas · 2026-08-24
- NAPE Audio Pretraining Achieves SOTA Without Decoders — kastnerkyle · 2026-08-24
- H3 excels at generating complex space scenes — SIR_NVAX_A_LOT · 2026-08-24