Tencent open-sources Prism: dynamic sparse attention for native 2K joint video-audio generation
linoy_tsaban · x · 2026-10-06
Tencent Hunyuan, with Fudan and Zhejiang University, open-sourced Prism on Hugging Face (MIT license), a video generation model natively trained for 2K joint video-audio generation.
Key ideas:
- Dynamic sparse attention replaces full attention: tokens are organized into spatiotemporal macro-zones whose attention structure adapts to local content
- Per-zone information structure is estimated via channel-wise variance of video features plus feature norms from audio-to-video cross-attention, then each zone gets a tailored block shape
- This counters the quadratic cost and diluted learning signals of full attention at high resolution, exploiting that cross-modal interactions concentrate around sound-producing regions
Preview clips look solid; paper: arXiv:2610.05416.
Related event: Tencent Hunyuan Open-Sources Prism for Native 2K Video-Audio Generation(3 posts)→
More from Multimodal
- inclusionAI open-sources Ming-Image-0.1-Design: 6B text-to-image model with transparent RGBA output — lmoroney · 2026-10-06
- Local Minimax H3 2K video generation on a 5090: 1984x1120 in 30 minutes, full prompt shared — Kooky-Mode3047 · 2026-10-06
- Wan2GP memory management update: 15s video now renders in 5 minutes, first-click OOM fixed — orangpelupa · 2026-10-06
- MIRRORSIDE: a behind-the-scenes film set that never existed, made with AI — StrategyMedium5907 · 2026-10-06
- Qwen-Image 2.1 edits look unfinished: user shares ComfyUI params seeking fixes — Suspicious_Aide2697 · 2026-10-06
- Creator hands repetitive workflow to Codex and GPT-6 Astra, keeps creative calls — socialwithaayan · 2026-10-06