Google's Active Video Understanding Cuts Gemini Token Use by 88% While Boosting Accuracy 7%
APPSO · wechat · 2026-09-03
Google has rolled out "Active Video Understanding" for Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite, letting the model autonomously decide which video segments to inspect, at what speed, and whether to rely on visuals, audio, or captions — replacing fixed frame sampling that misses fleeting actions.
- Benchmarks: Token consumption drops up to 88%, analysis cost up to 66%, with accuracy gains up to 7%; on LongVideoBench, Gemini 3.7 Flash sits on the accuracy-cost Pareto frontier.
- Availability: Live in Google AI Studio and Gemini Enterprise Agent Platform via the Gemini API (uploaded videos and YouTube), billed at standard token rates; coming to Gemini App Flash users and YouTube's AskYouTube in coming months.
- Use cases envisioned: Real-time sports tactics breakdown, semantic home-camera alerts, and ask-anything long-video search.
Related event: Google DeepMind Launches Agentic Video for Gemini: Up to 88% Fewer Tokens(28 posts)→
More from Models
- Heterogeneous prefill/decode isn't new: dev tried 70B prefill + 7B decode in 2023, 'entirely useless' — bingxu_ · 2026-09-03
- Using one model to audit three vendors' AI incident registry — from evidence QA to fixing the frontend — ValehartProject · 2026-09-03
- Ex-Cursor RL Lead Joins Meta's Reasoning Team as Muse Spark 1.3 Ships — ananyaku · 2026-09-03
- Devs warn using Gemini outside official surfaces can get your entire Google account banned — GlenBradley · 2026-09-03
- ChatGPT nails a backgammon dice probability problem at 47% odds — ivan_bezdomny · 2026-09-03
- H3 Acceleration Arena needs 1,700 more votes to crown the best turbo LoRA — Obvious_Set5239 · 2026-09-03