dRAE scales visual tokenization to 131,072 codes without codebook collapse

burny_tech · x · 2026-07-28

The paper introduces dRAE, a discrete representation autoencoder that uses hyper-spherical quantization instead of Euclidean codebook distances. The authors argue this avoids codebook collapse, better matches the geometry of visual representations, and allows much larger vocabularies.

According to the abstract, the method scales to 131,072 tokens, reaches 100% codebook utilization, simplifies training, and improves reconstruction, multimodal understanding, and image generation. The work positions angular routing as a better way to discretize high-dimensional visual features for language-model interfaces.

Original post →

More from Multimodal

Multimodal channel →