AutoGaze reduces visual tokens by 4x-100x for efficient 4K video understanding

Cohere · youtube · 2026-08-15

AutoGaze is a lightweight module designed to address spatial-temporal redundancy in multi-modal large language models (MLLMs) when processing long, high-resolution videos. Instead of processing every pixel equally, AutoGaze autoregressively selects a minimal set of multi-scale patches that can reconstruct the video within a specified error threshold before the data reaches the Vision Transformer (ViT) or MLLM.

Key Results:

The research is presented by Baifeng Shi, a researcher at Physical Intelligence and a PhD graduate from UC Berkeley (BAIR).

Original post →

More from Research

Research channel →