Gemini's 88% video token cut landed on 3.7 Flash, not the 3.8 everyone's talking about
Servola-Journal · reddit · 2026-09-03
Google shipped agentic video understanding to Gemini 3.7 Flash, 3.6 Flash and 3.5 Flash-Lite on September 1 — the day before the 3.8 Flash release everyone's discussing. The author argues the earlier rollout is the one that moves money.
The key change: instead of sampling every frame at a fixed rate, the model picks which segments to examine and whether to read frames, audio, or transcript. Google claims up to 88% fewer tokens and 66% lower cost at unchanged API pricing — turning camera-feed review from a budget line into a rounding error.
Caveat: the "up to 7% higher accuracy" claim doesn't specify test conditions, and a model that chooses what to skip should miss things a fixed scan catches. The author asks for tests against long footage with known ground truth.
More from Models
- llama.cpp adds support for NVIDIA's Nemotron-3-Puzzle-75B-A9B MoE model — pmttyji · 2026-09-03
- NVIDIA's 75B hybrid MoE Nemotron-3-Puzzle is now runnable locally in llama.cpp — jacek2023 · 2026-09-03
- Dev: GPT models lead at computer use, Anthropic far behind despite Fable's coding lead — CtrlAltDwayne · 2026-09-03
- GPT-6 spotted in branch name, rumored to focus on computer use capabilities — scaling01 · 2026-09-03
- GPT-5.6 mirrors your tone and swearing, while Claude stays stiff, user says — CtrlAltDwayne · 2026-09-03
- Miles Brundage: Fable 5.1 is far more token-hungry in real workflows than in evals — Miles_Brundage · 2026-09-03