Gemini agentic video understanding launches with a developer guide

osanseviero · x · 2026-09-02

Follow-up on Gemini's agentic video understanding: instead of spending 200k tokens on a long video, Gemini now uses tools to process video intelligently based on the prompt (analyzing transcription, adapting FPS, etc.), yielding 88% fewer tokens, lower latency, and higher accuracy. The author shares the developer guide and looks forward to unlocked use cases.

Related event: Google DeepMind rolls out Agentic Video Understanding for Gemini, cutting token use up to 88%(12 posts)→

Original post →

More from Models

Models channel →