LoHi: dense low-res sampling lifts long-video VLM accuracy 10.6 pts and cuts decoding latency 7x
jchoi2012 · hf · 2026-10-06
- The paper reframes efficient long-video understanding in VLMs: per-frame resolution can be traded for denser temporal coverage, and front-end decoding latency depends on candidate pool size rather than final token budget.
- Three findings: at matched token budgets, dense low-resolution sampling beats sparse native-resolution sampling; resolution-sensitive tasks benefit from a few selected high-res frames; front-end decoding dominates wall time for hour-long videos.
- LoHi is a training-free, single-pass framework combining a dense low-res video stream with sparse high-res image frames: LoHi-Anchor picks frames via codec I-frame metadata, LoHi-SemDiv uses query relevance and CLIP-feature diversity.
- Results: +10.6 pts average accuracy over native-resolution baseline and +5.2 pts over the strongest prior efficiency method across three benchmarks; up to 7x lower front-end decoding latency on hour-long videos.
More from Research
- Cornell Tech prof presents two pretraining-behavior papers at COLM, recruiting PhDs and postdocs — pratyushmaini · 2026-10-06
- Non-Expert Runs AI-Driven Interpretability Study Claiming Linear Representation of Directive Force in LLMs — LoudYogurtcloset7856 · 2026-10-06
- AC2 lets LLM RL train faster than GRPO by trusting critics on partial rollouts — stanfordnlp · 2026-10-06
- Stanford CS224N Winter 2026 posts free slides and assignments online — stanfordnlp · 2026-10-06
- Only right answers still teach chemistry: popular chemistry benchmark has shortcuts — tak3sh8 · 2026-10-06
- UT Austin math chair: OpenAI appears set to release ~400 AI-generated proofs at once — 141_1337 · 2026-10-06