Tencent Hunyuan's Prism: dynamic sparse attention speeds up 2K joint video-audio training 2.5x
Tencent-Hunyuan · hf · 2026-10-06
Tencent Hunyuan introduces Prism, a dynamic sparse attention framework for natively training joint video-audio generation models at 2K resolution.
- Problem: full attention's quadratic cost balloons at high resolution, and attention spreads over redundant tokens, diluting learning signals; existing sparse attention targets training-free speedups or ignores that cross-modal video-audio interactions concentrate around sound-producing regions.
- Method: Prism organizes tokens into spatiotemporal macro-zones, estimates each zone's information structure via channel-wise video feature variance and audio-to-video cross-attention norms, dynamically assigns tailored block shapes per zone (finer partitioning where visual content varies fast and audio-visual coupling is strong), and uses hybrid block selection for per-query sparsity.
- Results: 2.5x training speedup over full attention while surpassing it in generation quality.
More from Multimodal
- Martian showcases "Mars, Monumental" generated imagery — Kyrannio · 2026-10-06
- AI video made with Seedance 2.5 hailed as a masterpiece — SimplyAnnisa · 2026-10-06
- Dev resurfaces his 4-year-old UE5 AutoLOD tool for optimizing GenAI 3D game assets — rms80 · 2026-10-06
- AI 3D mesh generation tested: 1M+ polygons and terrible UVs, 'all terrible' — rms80 · 2026-10-06
- PoseForm: free browser tool to pose figures and export OpenPose/ControlNet references — Upstairs-Lead-2601 · 2026-10-06
- First Astana AI Film Festival draws 8,067 entries from 125 countries; $450K grand prize awarded — lmoroney · 2026-10-06