Kandinsky 6.0 Video: Open-Source Synchronized Audio-Video Models Under MIT

kandinskylab · hf · 2026-10-06

Kandinsky 6.0 Video: Foundation Models for Synced Audio-Video Generation

kandinskylab releases Kandinsky 6.0 Video, a family of diffusion foundation models for synchronized text-to-audio-video (T2AV) and image-to-audio-video (I2AV) generation, in Lite (3B) and Pro (29B) variants. They generate 5-second clips with synchronized 44 kHz audio including lip-sync; a built-in super-resolution model upscales output to Full-HD.

Technical highlights:

Evaluation: in human side-by-side comparisons, 6.0 Video Pro clearly beats Kandinsky 5.0 Video Pro and stays competitive with leading audio-video models, especially in speech quality. Code, checkpoints, and diffusers integration are released under MIT license.

Original post →

More from Multimodal

Multimodal channel →