Google's Active Video Understanding Cuts Gemini Token Use by 88% While Boosting Accuracy 7%

APPSO · wechat · 2026-09-03

Google has rolled out "Active Video Understanding" for Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite, letting the model autonomously decide which video segments to inspect, at what speed, and whether to rely on visuals, audio, or captions — replacing fixed frame sampling that misses fleeting actions.

Related event: Google DeepMind Launches Agentic Video for Gemini: Up to 88% Fewer Tokens(28 posts)→

Original post →

More from Models

Models channel →