StableVQ: three lightweight fixes for stable vector-quantized tokenizer training

Kwai-Kolors · hf · 2026-09-23

Kwai-Kolors (Kuaishou) released StableVQ, a systematic take on the underexplored problem of training stability in discrete visual tokenizers.

Diagnosis

Three fixes

Results: Built on shared-projection codebooks, StableVQ is lightweight and adds no learnable parameters; on ImageNet it consistently improves training stability, codebook utilization, and reconstruction quality across codebook sizes and initializations.

Original post →

More from Multimodal

Multimodal channel →