MIT releases open-weight Kandinsky video+audio model in 6 variants with Diffusers day-0 support

RisingSayak · x · 2026-10-06

MIT has released Kandinsky, an open-weight video model with joint audio generation. It ships in 6 variants balancing speed, memory, and quality, plus two super-resolution checkpoints. Hugging Face Diffusers offers day-0 integration, so the model is usable on release day.

Related event: Kandinsky 6.0 Video Open-Sources Audio-Video Generation Under MIT License(4 posts)→

Original post →

More from Multimodal

Multimodal channel →