AutoGaze Slashes Visual Tokens by 100x for High-Res Long Video Understanding
Cohere_Labs · x · 2026-08-06
The Cohere Labs Open Science Community announced an upcoming technical webinar focused on AutoGaze, presented by Baifeng Shi, a researcher at Physical Intelligence and UC Berkeley PhD.
The research addresses the computational bottlenecks and spatiotemporal redundancy when Multi-modal Large Language Models (MLLMs) process long, high-resolution videos.
- Core Tech: AutoGaze is a lightweight module that uses next-token prediction and reinforcement learning to autoregressively select a minimal set of multi-scale patches before they are processed by a ViT or MLLM.
- Performance: It reduces visual tokens by 4x to 100x and accelerates ViTs and MLLMs by up to 19x, enabling MLLMs to scale to 1,000-frame 4K-resolution videos.
- Benchmarks: Achieves 67.0% on VideoMME. The team also introduces HLVid, the first high-resolution, long-form video QA benchmark with 5-minute 4K videos, where the AutoGaze-scaled MLLM outperforms the baseline by 10.1% and the previous best MLLM by 4.5%.
More from Research
- ICML 2026 Oral Papers Show Poor Reproducibility; Researcher Vows Not to Hire Those with <50% Reproducible Papers — ChenhaoTan · 2026-08-06
- The Secret to Good Data: Punish Lazy Automation and Embrace Organic Diversity — tokenbender · 2026-08-06
- ECCV 2026 Paper: Privacy Leakage in Scene Coordinate Regression Models — CSProfKGD · 2026-08-06
- CMU Combines Theorem Provers and Neural AI to Reshape Mathematics — AkariAsai · 2026-08-06
- Flux 3's Core Tech 'Self-Flow' Breaks External Alignment Bottleneck — linoy_tsaban · 2026-08-06
- ARC Benchmark Criticized for Ignoring Tool Use, Missing the Actual AI Trend — scaling01 · 2026-08-06