Fudan, Tencent Hunyuan and Zhejiang U. release Prism: sparse attention trains video+audio models 2.5x faster

lmoroney · x · 2026-10-07

Prism, from Fudan University, Tencent Hunyuan, and Zhejiang University, is a sparse attention method for natively training joint video-and-audio models at 720p, 1080p, and 2K (2560x1440). Code, paper, and two preview checkpoints are out.

Related event: Tencent Hunyuan Open-Sources Prism for 2K Audio-Video Generation(4 posts)→

Original post →

More from Multimodal

Multimodal channel →