Study: frame selection dominates long-video MLLM accuracy; spatial compression is nearly free if reinvested

Prakhar Khatri · hf · 2026-09-04

A controlled study of visual-token allocation in long-video MLLMs finds that frame selection dominates accuracy, while spatial compression is nearly free when the saved tokens are reinvested into more frames. The authors also propose a unified harness for fairly comparing frame selectors.

Original post →

More from Multimodal

Multimodal channel →