dRAE scales visual tokenization to 131,072 codes without codebook collapse
burny_tech · x · 2026-07-28
The paper introduces dRAE, a discrete representation autoencoder that uses hyper-spherical quantization instead of Euclidean codebook distances. The authors argue this avoids codebook collapse, better matches the geometry of visual representations, and allows much larger vocabularies.
According to the abstract, the method scales to 131,072 tokens, reaches 100% codebook utilization, simplifies training, and improves reconstruction, multimodal understanding, and image generation. The work positions angular routing as a better way to discretize high-dimensional visual features for language-model interfaces.
More from Multimodal
- Netflix’s ID-V2V preserves identity while restyling videos from one source clip — netflix · 2026-07-28
- New paper maps compute-optimal scaling laws for native multimodal pre-training — burny_tech · 2026-07-28
- New ComfyUI node converts audio into MIDI for music workflows — MuziqueComfyUI · 2026-07-28
- Boogu-Image-0.1 says 2.08 billion images and $400,000 were enough to reach open-source SOTA — 机器之心 · 2026-07-28
- Exploring ComfyUI Basics: Why Separate Checkpoint and KSampler in Workflows? — DavidThi303 · 2026-07-28
- Running Z Image Turbo on RTX 4060: How to Break Through the Quality Ceiling? — Dangerous_Ring_435 · 2026-07-28