Image-to-video model Prism with joint video-audio generation trends on Hugging Face, MIT licensed
FrancisRing · hf · 2026-10-07
FrancisRing's Prism is trending on Hugging Face: an image-to-video video diffusion transformer featuring joint video-audio generation, sparse attention, and high-resolution output.
It ships with diffusers and safetensors support, references arxiv:2610.05416, and is released under the MIT license.
More from Multimodal
- Voice Agents Live or Die on 'Sounding Right': Turbo Shifts Tone With User Emotion — SucceededMind · 2026-10-07
- Magnific Original Series The Chronicles of Bone drops Chapter Six, made entirely with AI tools — Kavanthekid · 2026-10-07
- Hedra Lands in ChatGPT: Attach One Product Photo, Get a Full Commercial Ad — henloitsjoyce · 2026-10-07
- Marc Andreessen boosts AI film contest SLOPTOBERFEST grand prize to $25,000 — zealcaiden · 2026-10-07
- Image generation pricing leak: $0.05 per 2K image, $0.076 per 4K — op7418 · 2026-10-07
- Live human votes plugged into Flow-GRPO to stop image models gaming reward models — lmoroney · 2026-10-07