OmniScope: Training-Free Token Compression for Omnimodal LLMs

Jinsen Su · hf · 2026-08-01

OmniScope introduces a training-free token compression framework for omnimodal large language models. It addresses the issue where existing methods discard answer-critical cues due to cross-modal salience mismatches (e.g., audio and video relevance peaking at different moments).

Core Mechanisms:

Results: Across 4 audio-video benchmarks and 2 Qwen2.5-Omni model scales, OmniScope achieves the best average accuracy. At 25% token retention, it delivers up to a 3.53x prefill speedup and over 15% GPU memory reduction, with only a 0.35-point drop in average accuracy.

Original post →

More from Infra

Infra channel →