Kandinsky Lab Releases KVAE: Family of Tokenizers for Multimodal Models

kandinskylab · hf · 2026-08-08

Kandinsky Lab released the KVAE series of tokenizers, designed for text-conditioned generation across audio, image, and video modalities:

The team claims that KVAE matches or surpasses frontier open-source tokenizers (e.g., Wan-2.2, FLUX.2, StableAudio) on both objective and subjective reconstruction and generation metrics. They also shared training details, model selection methods, and ablations, with code fully open-sourced.

Original post →

More from Multimodal

Multimodal channel →