Gemini adds agentic video understanding, cutting token usage by 88%

osanseviero · x · 2026-09-02

Gemini now supports agentic video understanding: instead of spending 200k tokens processing a long video, it uses tools to process video intelligently based on the prompt (analyzing transcription, adapting FPS, etc.). This leads to 88% fewer tokens, lower latency, and higher accuracy. A developer guide is available.

Related event: Google DeepMind rolls out Agentic Video Understanding for Gemini, cutting token use up to 88%(12 posts)→

Original post →

More from Models

Models channel →