Kandinsky Lab Releases KVAE: Family of Tokenizers for Multimodal Models
kandinskylab · hf · 2026-08-08
Kandinsky Lab released the KVAE series of tokenizers, designed for text-conditioned generation across audio, image, and video modalities:
- KVAE-Audio: A continuous full-band 48 kHz tokenizer with a 50 Hz latent of 64 channels.
- KVAE-3D: Two causal video tokenizers for 4x16x16 and 4x8x8 compression.
- KVAE-2D: An image model compressing input by a factor of 8 with 32 channels.
The team claims that KVAE matches or surpasses frontier open-source tokenizers (e.g., Wan-2.2, FLUX.2, StableAudio) on both objective and subjective reconstruction and generation metrics. They also shared training details, model selection methods, and ablations, with code fully open-sourced.
More from Multimodal
- fal Launches Agent: A Creative Assistant Integrating Image, Video, and 3D Models — gorkem · 2026-08-08
- Training Image Editing Models Without Human Labels via Video Deltas — haremlifegame · 2026-08-08
- Ostris Teases Upcoming LoRA Weights for MiniMax H3 Model — krigeta1 · 2026-08-08
- Seedance 2.5 video-to-video test: 720p results show promise — mrjonfinger · 2026-08-08
- DeepSeek and MiniMax Collaborate: Hermes Agent Autonomously Completes Game Design — NousResearch · 2026-08-08
- Creator Shares Ambitious Sci-Fi AI Short Film 'Second Earth' — BradClarkAI · 2026-08-08