Braco compresses visual tokens 144x at 95.2% accuracy with ~36% end-to-end speedup

Rui Zhong · hf · 2026-09-30

The paper revisits extreme visual-token compression in VLMs through a token-parameterization lens, separating basis transformation and structured truncation (compressibility) from coordinate organization (learnability and cross-modal alignment).

Original post →

More from Research

Research channel →