Kandinsky 6: MIT open-weight video+audio model with 6 variants, day-0 Diffusers support
RisingSayak · x · 2026-10-06
Kandinsky Lab released Kandinsky 6, an open-weight (MIT) video generation family with day-0 integration in Hugging Face Diffusers.
Key details:
- The main model is a single multimodal diffusion transformer that generates video and synchronized audio from text or a reference image: video and audio latents are jointly denoised through fused blocks with cross-attention, each conditioned on its own Qwen2.5-VL text branch and a CLIP pooled embedding.
- A separate super-resolution model upscales output tile-by-tile in the latent space of a causal 3D K-VAE.
- Six variants balance speed, memory, and quality; official checkpoints include flow-matching and distilled versions plus two super-res checkpoints.
Related event: Kandinsky 6.0 Video Open-Sources Audio-Video Generation Under MIT License(4 posts)→
More from Multimodal
- Runway announces World Runner: generative worlds on a pocket Game Boy-style console — c_valenzuelab · 2026-10-06
- Gothic vampire short film made with Seedance 2.5 on Runwayml — azed_ai · 2026-10-06
- AI-generated superheroes just want chai and biscuits, not saving the world — umesh_ai · 2026-10-06
- YarnGPT quietly ships Pidgin voice translation, more languages coming — saheedniyi_02 · 2026-10-06
- Doubao and Qwen voice models clone timbre from ~10s of audio, cheap enough for agents — AlchainHust · 2026-10-06
- Kling 4.0 Flash dialogue quality impresses: characters now act with pauses and eye contact — LudovicCreator · 2026-10-06