Tencent open-sources Prism: dynamic sparse attention for native 2K joint video-audio generation

linoy_tsaban · x · 2026-10-06

Tencent Hunyuan, with Fudan and Zhejiang University, open-sourced Prism on Hugging Face (MIT license), a video generation model natively trained for 2K joint video-audio generation.

Key ideas:

Preview clips look solid; paper: arXiv:2610.05416.

Related event: Tencent Hunyuan Open-Sources Prism for Native 2K Video-Audio Generation(3 posts)→

Original post →

More from Multimodal

Multimodal channel →