KATok: adaptive video tokenizer drops uninformative tokens for compact representation

kakaocorp · hf · 2026-09-01

KATok is a transformer-based adaptive video tokenizer from Kakao that selectively drops uninformative tokens to achieve data-dependent compression while preserving spatial consistency for diffusion-based video generation.

Unlike fixed-ratio schemes, it yields more compact video representations for downstream diffusion generation.

Original post →

More from Multimodal

Multimodal channel →