KATok: adaptive video tokenizer drops uninformative tokens for compact representation
kakaocorp · hf · 2026-09-01
KATok is a transformer-based adaptive video tokenizer from Kakao that selectively drops uninformative tokens to achieve data-dependent compression while preserving spatial consistency for diffusion-based video generation.
Unlike fixed-ratio schemes, it yields more compact video representations for downstream diffusion generation.
More from Multimodal
- Krea2-SDA LoRA fixes lack of variety in Krea2 Turbo generations — Gradio · 2026-09-01
- Chat-Edit-3D++: LLM-Driven Conversational Editing of 3D and 4D Scenes — stepfun-ai · 2026-09-01
- CPU-only video workflow for blurring background crowd — Exotic_Accountant565 · 2026-09-01
- User Questions Minimax H3 Hybrid Models: Weak Reference Adherence — Tablaski · 2026-09-01
- ABot-Recon lands on Hugging Face: turns dashcam clips into 3D scans in seconds — petewoodbridge · 2026-09-01
- DICS: Exploring Data Intrinsic Consistency for Visual Instruction Selection — SpatialAxiom · 2026-09-01