QuadTok: quadtree visual tokenizer saves 10% tokens, hits 2.08 gFID on ImageNet 256
Yucheng Mao · hf · 2026-10-08
QuadTok is a hierarchical quadtree visual tokenizer for autoregressive image generation, bridging 2D spatial binding and 1D sequence flexibility.
Highlights:
- Dynamically allocates capacity: fine regions get more tokens, homogeneous regions stay coarse.
- Saves 10% tokens vs a fixed 256-token grid on ImageNet and 9% zero-shot on COCO, with comparable reconstruction fidelity.
- The tree structure naturally enables autoregressive generation: a 947M GPT-style model conditioned on the quadtree topology reaches 2.08 gFID on ImageNet 256×256.
- Spatial correlation preserved by the quadtree enables zero-shot spatially controlled generation.
Code is available on GitHub.
More from Multimodal
- Midjourney prompt shares Kirlian-style glowing aura portraits with --v 8.2 — tisch_eins · 2026-10-08
- VELA 1.0 speeds up MiniMax H3 video gen: 0.8MP render drops from 2:53 to 2:13 — NoMouse9610 · 2026-10-08
- InSpatio-World 1.5 open-sources real-time 4D world model with multi-input support — liuziwei7 · 2026-10-08
- Opus 5.5 drives Blender+Suno via MCP, delivers sync'd 20s masterpiece in 3.5 hours with zero keyframes — sidahuj · 2026-10-08
- User shares strikingly realistic GPT-6 fantasy hunter character portrait — TheUltraBased · 2026-10-08
- LLM-coded Blender beats video models at controllable video generation — jyangballin · 2026-10-08