V-CoLA: training-free vision token compression keeps 99.5% performance at half the tokens
ATH-MaaS · hf · 2026-10-09
ATH-MaaS released V-CoLA, a training-free vision token compression framework designed for linear-attention hybrid VLM architectures (e.g., Qwen3.5).
Key points:
- Prior compression methods built for softmax attention degrade noticeably on hybrid linear-attention architectures.
- V-CoLA introduces a uniqueness-aware importance criterion for selecting critical vision tokens plus an adaptive token merging strategy, with implementations compatible with chunk-wise parallelism of linear attention.
- Results: 99.5% of original performance with 50% of vision tokens, over 88% with just 12.5%, and prefill speedups of 1.86x to 6.15x across benchmarks.
More from Models
- Qwen3.8-Max, Flash and Wan3.0 free for a week on GMI Cloud with higher rate limits — Alibaba_Qwen · 2026-10-09
- ByteDance Seed team spotted DeepSeek performance drift, Zhihu explainer goes viral — teortaxesTex · 2026-10-09
- Chollet: far transfer has no evidence in humans, and LLMs/LRMs may be no different — burny_tech · 2026-10-09
- Google AI Overviews hallucinates while user searches for a meme — prajdabre · 2026-10-09
- InfiLoop residuals let a 7M looped model improve past 20,000 test-time steps, 97.9% on Sudoku — Pengxiang Li · 2026-10-09
- ChatGPT Pro users report usage cut from 20x to 10x, but free weekly resets restore it — vuonghtt · 2026-10-09