AutoGaze reduces visual tokens by 4x-100x for efficient 4K video understanding
Cohere · youtube · 2026-08-15
AutoGaze is a lightweight module designed to address spatial-temporal redundancy in multi-modal large language models (MLLMs) when processing long, high-resolution videos. Instead of processing every pixel equally, AutoGaze autoregressively selects a minimal set of multi-scale patches that can reconstruct the video within a specified error threshold before the data reaches the Vision Transformer (ViT) or MLLM.
Key Results:
- Efficiency: Reduces visual tokens by 4x-100x and accelerates ViTs and MLLMs by up to 19x.
- Scalability: Enables MLLMs to process 1,000-frame 4K-resolution videos.
- Performance: Achieves 67.0% on the VideoMME benchmark. On the newly introduced HLVid benchmark (5-minute 4K video QA), it improves the baseline by 10.1% and outperforms the previous best MLLM by 4.5%.
The research is presented by Baifeng Shi, a researcher at Physical Intelligence and a PhD graduate from UC Berkeley (BAIR).
More from Research
- MONA: Myopic Optimization Mitigates Multi-step Reward Hacking in RL — sebkrier · 2026-08-15
- Sébastien Bubeck's book on Convex Optimization available on ChapterPal — burkov · 2026-08-15
- Authors unpack viral 100-page paper on reasoning heist and model distillation — burny_tech · 2026-08-15
- CRISPR screen vs aging atlas: phase imaging wins per dollar; tissue aging is supracellular — anshulkundaje · 2026-08-15
- RoMaV2: Harder, Better, Faster, Denser Feature Matching Model — tom_doerr · 2026-08-15
- Daily AI Reading: Speeding up generative UI and multi-agent coordination patterns — rseroter · 2026-08-15