Kandinsky 6.0 Video: Open-Source Synchronized Audio-Video Models Under MIT
kandinskylab · hf · 2026-10-06
Kandinsky 6.0 Video: Foundation Models for Synced Audio-Video Generation
kandinskylab releases Kandinsky 6.0 Video, a family of diffusion foundation models for synchronized text-to-audio-video (T2AV) and image-to-audio-video (I2AV) generation, in Lite (3B) and Pro (29B) variants. They generate 5-second clips with synchronized 44 kHz audio including lip-sync; a built-in super-resolution model upscales output to Full-HD.
Technical highlights:
- Dual-stream CrossDiT architecture links a pretrained video stream and a newly trained audio stream via bidirectional cross-attention for temporal and semantic alignment;
- Continuous pretraining: train the audio stream from scratch on audio-only corpora, then jointly train both streams on paired audio-video data preserving unimodal fidelity, followed by SFT, RL-based post-training, and distillation.
Evaluation: in human side-by-side comparisons, 6.0 Video Pro clearly beats Kandinsky 5.0 Video Pro and stays competitive with leading audio-video models, especially in speech quality. Code, checkpoints, and diffusers integration are released under MIT license.
More from Multimodal
- MIRRORSIDE: a behind-the-scenes film set that never existed, made with AI — StrategyMedium5907 · 2026-10-06
- Qwen-Image 2.1 edits look unfinished: user shares ComfyUI params seeking fixes — Suspicious_Aide2697 · 2026-10-06
- Creator hands repetitive workflow to Codex and GPT-6 Astra, keeps creative calls — socialwithaayan · 2026-10-06
- Midjourney style code sref 705714994 yields whimsical brush-texture characters — michaelrabone · 2026-10-06
- AI-Created Crowd Looks Indistinguishable From Real People, Creator Warns — minchoi · 2026-10-06
- Runway CEO invokes Jevons paradox: coding agents are expanding software demand, video is next — c_valenzuelab · 2026-10-06