OmniDelta trims OmniLLM audio-video tokens with skill-driven budget allocation

Haoyang Huang · hf · 2026-07-29

OmniDelta reallocates token budgets for audio-video compression in OmniLLMs

OmniDelta is a training-free, skill-driven framework for token compression in omni-modal LLMs. The paper argues that existing pruning methods focus too much on selecting tokens under a fixed budget, while the harder problem is how to allocate that budget across modalities and within each modality.

What it does

Results

Original post →

More from Multimodal

Multimodal channel →