OmniScope: Training-Free Token Compression for Omnimodal LLMs
Jinsen Su · hf · 2026-08-01
OmniScope introduces a training-free token compression framework for omnimodal large language models. It addresses the issue where existing methods discard answer-critical cues due to cross-modal salience mismatches (e.g., audio and video relevance peaking at different moments).
Core Mechanisms:
- Modality Decoupling: Uses the query as a shared semantic anchor but estimates relevance and allocates token budgets separately for audio and video.
- Visual Pruning: Employs an anchor-delta strategy to preserve both global context and temporal changes.
- Audio Merging: Merges audio tokens within each second to reduce redundancy while maintaining temporal continuity.
Results: Across 4 audio-video benchmarks and 2 Qwen2.5-Omni model scales, OmniScope achieves the best average accuracy. At 25% token retention, it delivers up to a 3.53x prefill speedup and over 15% GPU memory reduction, with only a 0.35-point drop in average accuracy.
More from Infra
- Benchmark: NInfer Boosts Qwen3.6 Prefill Speed Over 2x vs llama.cpp — tat_tvam_asshole · 2026-08-01
- Jensen Huang Slams GPU-to-Nuke Analogy: Everyone Should Have AI — rohanpaul_ai · 2026-08-01
- DeepSeek-V4-Flash Inference Blocked: vLLM Lacks Support for New confidence_head — teortaxesTex · 2026-08-01
- AMD Open-Sources Mi455/CDNA5 GPU Instruction Set Architecture — AnushElangovan · 2026-08-01
- Moonshot's MoonEP Acknowledges Alibaba's AcclEP, A Library with Zero Public Source — giffmana · 2026-08-01
- AMD iGPU LLM Benchmark: MoE Architectures Crush Dense Models on Edge Inference — tabletuser_blogspot · 2026-08-01