Fudan, Tencent Hunyuan and Zhejiang U. release Prism: sparse attention trains video+audio models 2.5x faster
lmoroney · x · 2026-10-07
Prism, from Fudan University, Tencent Hunyuan, and Zhejiang University, is a sparse attention method for natively training joint video-and-audio models at 720p, 1080p, and 2K (2560x1440). Code, paper, and two preview checkpoints are out.
- Method: splits the token sequence into zones, each with its own attention block shape—finer blocks where the picture changes quickly or sound is strongly tied to on-screen content.
- Results: authors report 2.5x training speedup over full attention, with improved quality.
- Hardware: 720p inference fits on one 80 GB GPU; 1080p and 2K need four or more; native 720p training starts at 32 such GPUs.
- Getting started: rent a single 80 GB card, start with 720p, and gauge outputs before scaling up.
Related event: Tencent Hunyuan Open-Sources Prism for 2K Audio-Video Generation(4 posts)→
More from Multimodal
- AI video contest winner reimagines Louis XIV's court colonizing Mars — Loo_Atreides · 2026-10-07
- AI comic magazine TischLog #48 explores reversed morality, built with Midjourney V8.2 — tisch_eins · 2026-10-07
- MEND: RL for flow models via proximal velocity matching beats Flow-GRPO in 100 vs ~4k updates — UTEXAS · 2026-10-07
- JLD: perceptual distance from a frozen encoder's Jacobian, fitted in 35s from 100 images, beats LPIPS and DISTS — Shreshth Saini · 2026-10-07
- Qwen-Image Multiple-Angles LoRA released on Hugging Face — Illustrious_Row_9971 · 2026-10-07
- Dev ships Vulkan-only local generative art app with no telemetry or cloud — ogimaru · 2026-10-07