Image-to-video model Prism with joint video-audio generation trends on Hugging Face, MIT licensed

FrancisRing · hf · 2026-10-07

FrancisRing's Prism is trending on Hugging Face: an image-to-video video diffusion transformer featuring joint video-audio generation, sparse attention, and high-resolution output.

It ships with diffusers and safetensors support, references arxiv:2610.05416, and is released under the MIT license.

Original post →

More from Multimodal

Multimodal channel →