CAS Introduces GMC: 90% Visual Token Compression with No Performance Drop

量子位 · wechat · 2026-08-12

Vision-Language Models (VLMs) often struggle with high computational and memory costs due to massive visual tokens from high-res images. While traditional Top-K pruning causes detail loss and hallucinations, the Zidong Taichu team at CAS introduced GMC (Grounded Message Coreset Pruning). This training-free, two-stage method uses complementary evidence selection and population transport to achieve high-fidelity compression.

Experiments on models like Qwen2.5-VL show that retaining only 10% of visual tokens preserves over 99% of the original performance, occasionally exceeding it due to noise reduction. The method also achieves a 1.25x end-to-end inference speedup and significantly reduces KVCache usage in long-document scenarios, effectively solving the trade-off between compression and accuracy.

Original post →

More from Models

Models channel →