Gemini API launches Agentic Video: up to 88% fewer tokens and better long-video reasoning

davidstutz92 · x · 2026-09-04

Google DeepMind's Logan Kilpatrick introduced Agentic Video in the Gemini API, a new way to process long videos that cuts token consumption by up to 88% while increasing quality, controllable per video and available on the newest models like 3.7 Flash. A third-party retweet reports a +5.3pp accuracy jump on the Minerva long-video reasoning benchmark at only 42% of the token cost. The benchmark comes from DeepMind's open-sourced Minerva Dataset Collection, including Minerva-cultural: 2,200 human-crafted QA pairs across 540 culturally rich videos in 18 locales for testing multilingual, multicultural long-video reasoning in Video-LLMs.

Original post →

More from Models

Models channel →