Google's KeyRec Achieves Best Long-Video VLM Results With Just 10% of Visual Token Budget
google · hf · 2026-10-06
Google proposes KeyRec, a training-free framework for bounded visual memory in streaming and long-video understanding, tackling the explosion of dense visual tokens as video length grows.
Design
- Query-agnostic writing: a visual cache preserves fine-grained recent observations while historical evidence is organized into a structured event bank, updated online via add-merge-evict based on novelty.
- On question arrival, a text-only router adaptively splits a fixed readout budget between recent and event memory, without reprocessing historical frames.
- Operates on model-facing visual embeddings, supporting both encoder-projector VLMs and the encoder/projector-free NEO-ov architecture.
Results
- Across 4 streaming/long-video benchmarks and 3 VLM backbones: best compressed performance in 13 of 15 settings using only 10% of the dense visual-token budget.
- Outperforms the strongest compressed baseline by 2.21-18.37 points on real-time questions; best in 5 of 6 long-video settings and every NEO-ov 2B setting.
More from Multimodal
- One Universal Prompt Turns Any Product Photo Into a Cinematic Ad Across ChatGPT, Nano-Banana and Seedream — aziz4ai · 2026-10-06
- Suno Viral Creator Dream Relic Drops New Album 'Lost In a Dream' — suno · 2026-10-06
- First try with HyperFrames desktop app: one-prompt "vibe editing" looks surprisingly good — toolstelegraph · 2026-10-06
- AssemblyAI's Universal 3.6 Pro cuts voice-agent transcription errors 45%, adds 14 languages — AssemblyAI · 2026-10-06
- From Will Smith eating spaghetti like an alien to him fighting spaghetti in 3 years — Calm_Cartographer324 · 2026-10-06
- Open-source fibo-scene-analyzer outputs structured captions, bboxes and poses in one model — linoy_tsaban · 2026-10-06