Gemini's 88% video token cut landed on 3.7 Flash, not the 3.8 everyone's talking about

Servola-Journal · reddit · 2026-09-03

Google shipped agentic video understanding to Gemini 3.7 Flash, 3.6 Flash and 3.5 Flash-Lite on September 1 — the day before the 3.8 Flash release everyone's discussing. The author argues the earlier rollout is the one that moves money.

The key change: instead of sampling every frame at a fixed rate, the model picks which segments to examine and whether to read frames, audio, or transcript. Google claims up to 88% fewer tokens and 66% lower cost at unchanged API pricing — turning camera-feed review from a budget line into a rounding error.

Caveat: the "up to 7% higher accuracy" claim doesn't specify test conditions, and a model that chooses what to skip should miss things a fixed scan catches. The author asks for tests against long footage with known ground truth.

Related event: Google DeepMind Launches Agentic Video for Gemini, Cutting Token Use by Up to 88%(29 posts)→

Original post →

More from Models

Models channel →