Kandinsky 6.0 Video open-sourced under MIT, vLLM adds day-0 inference support

vllm_project · x · 2026-10-06

Kandinsky 6.0 Video is open-sourced under MIT, offering Lite (3B) and Pro (29B) diffusion models that generate 5-second video clips with synchronized 44kHz audio and lip-sync, in both text-to-audio-video (T2AV) and image-to-audio-video (TI2AV) modes. A plugin-in super-resolution model raises output to Full-HD (1920×1080).

vLLM-Omni announced day-0 support, so the models can run inference on vLLM from launch day. The repo also ships a ComfyUI plugin and quick-start scripts; requires an NVIDIA GPU and Python 3.13+, with the Pro Distill 5s weights downloaded automatically on first run.

Related event: Kandinsky 6.0 Video Open-Sources Audio-Video Generation Under MIT License(4 posts)→

Original post →

More from Multimodal

Multimodal channel →