CAS Introduces GMC: 90% Visual Token Compression with No Performance Drop
量子位 · wechat · 2026-08-12
Vision-Language Models (VLMs) often struggle with high computational and memory costs due to massive visual tokens from high-res images. While traditional Top-K pruning causes detail loss and hallucinations, the Zidong Taichu team at CAS introduced GMC (Grounded Message Coreset Pruning). This training-free, two-stage method uses complementary evidence selection and population transport to achieve high-fidelity compression.
Experiments on models like Qwen2.5-VL show that retaining only 10% of visual tokens preserves over 99% of the original performance, occasionally exceeding it due to noise reduction. The method also achieves a 1.25x end-to-end inference speedup and significantly reduces KVCache usage in long-document scenarios, effectively solving the trade-off between compression and accuracy.
More from Models
- Report: Ilya Sutskever's SSI Pivots to Test-Time Training for New Reasoning Engine — iruletheworldmo · 2026-08-12
- Liquid Crow Released: Squeezing Physical Cognition into a 450M-Parameter Micro-World Model — helloiamleonie · 2026-08-12
- Meta vs NVIDIA 30B Agent Models: Local Execution vs Cloud Routing — eyishazyer · 2026-08-12
- Too RAM-Hungry? Devs Discuss Best SLMs to Run Locally on 16GB Machines — elie2222 · 2026-08-12
- Paper Decodes Hidden Reasoning Traces from Claude and GPT — Zealousideal_Sort74 · 2026-08-12
- Developer Warning: Claude Opus and Fable Generate Code with Severe Security Flaws — OwariDa · 2026-08-12