Study: frame selection dominates long-video MLLM accuracy; spatial compression is nearly free if reinvested
Prakhar Khatri · hf · 2026-09-04
A controlled study of visual-token allocation in long-video MLLMs finds that frame selection dominates accuracy, while spatial compression is nearly free when the saved tokens are reinvested into more frames. The authors also propose a unified harness for fairly comparing frame selectors.
More from Multimodal
- Omni 1.1 adds video references: 3-second clips ensure character consistency in AI video — fofrAI · 2026-09-04
- AI-made Resident Evil x Psycho horror parody, part one — No-Invite8044 · 2026-09-04
- MiniMax H3 ecosystem boom: cinematic LoRAs, 360° video, and one-take chaining tools — optimisticalish · 2026-09-04
- Gemini-based agentic video understanding experiment goes viral, repo coming soon — MarioLucic_ · 2026-09-04
- Hands-on demo of Midjourney Edit shared on X — sergeantsref · 2026-09-04
- DeepLearning.AI and Google launch free course on agents for image and video generation — DeepLearningAI · 2026-09-04