Gemini agentic video understanding: 88% fewer tokens, 66% lower cost, ~7% higher accuracy

_philschmid · x · 2026-09-02

Phil Schmid breaks down Gemini's agentic video understanding: the model iteratively navigates video timelines, decides what to watch, picks frame rates (0.1 or 10 FPS), and chooses whether it needs speech transcripts, audio, or visual frames. Long videos get up to 88% fewer tokens, 66% lower costs, and 7% higher accuracy on benchmarks.

How it works:

Set processing="agentic" on video to enable; keep static for videos under 2 minutes. Available now for Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash Lite.

Original post →

More from coding & agent

coding & agent channel →