Braco compresses visual tokens 144x at 95.2% accuracy with ~36% end-to-end speedup
Rui Zhong · hf · 2026-09-30
The paper revisits extreme visual-token compression in VLMs through a token-parameterization lens, separating basis transformation and structured truncation (compressibility) from coordinate organization (learnability and cross-modal alignment).
- The two coupled objectives are formalized as unified functionals
- This yields Braco, a lightweight four-step coder combining transform-basis truncation, input-independent basis-coordinate embeddings, budget-dependent orthogonal re-parameterization, and residual spatial tokens from lightweight pooling
- Forms the favorable accuracy-efficiency frontier at 23x–64x compression and stays competitive at 144x: 95.2% accuracy with 84.2%–86.7% fewer prefill FLOPs than the uncompressed upper bound
- Versus prior methods: matches or improves accuracy with up to 36% end-to-end speedup and 16.6x/78.8x lower compressor latency/FLOPs
More from Research
- SCATR: a lightweight calibrated scorer finds the right LLM answer cheaply — pliang279 · 2026-09-30
- Simons Institute Report Lays Out 12 Actions for TCS Community to Respond to AI Progress — jasondeanlee · 2026-09-30
- Dev predicts a universal DSL for video data will enable highly controllable world-simulation diffusion models — zeeshanp_ · 2026-09-30
- arXiv tops 3.19 million submissions as monthly paper volume keeps climbing — _reachsumit · 2026-09-30
- Hugging Face ships tokenizers v1, often tens of times faster than v0.23 — ariG23498 · 2026-09-30
- EMNLP oral paper: RL helps models traverse parametric knowledge inaccessible after instruction tuning — niloofar_mire · 2026-09-30