Gemini API launches Agentic Video: up to 88% fewer tokens and better long-video reasoning
davidstutz92 · x · 2026-09-04
Google DeepMind's Logan Kilpatrick introduced Agentic Video in the Gemini API, a new way to process long videos that cuts token consumption by up to 88% while increasing quality, controllable per video and available on the newest models like 3.7 Flash. A third-party retweet reports a +5.3pp accuracy jump on the Minerva long-video reasoning benchmark at only 42% of the token cost. The benchmark comes from DeepMind's open-sourced Minerva Dataset Collection, including Minerva-cultural: 2,200 human-crafted QA pairs across 540 culturally rich videos in 18 locales for testing multilingual, multicultural long-video reasoning in Video-LLMs.
More from Models
- Gemini 3.8 Flash bug fixed: Google AI Mode now shows far more source links — gaganghotra_ · 2026-09-04
- Critics warn OpenAI's GPT-6 Astra reasons opaquely, gutting CoT monitoring safety — GaryMarcus · 2026-09-04
- ChatGPT adds writing-style matching from connected apps, analytics, and a Yubikey deal tied to Daybreak access — btibor91 · 2026-09-04
- GPT-6 reportedly launches as Tesla starts public rides in steering-free Cybercab — Dr_Singularity · 2026-09-04
- Astra early-access users' similar blender demos look coordinated, with no practical examples shown — jdjohnson · 2026-09-04
- Researcher teases dynamic composite eval index as "evals run on Twitter vibes" — evijit · 2026-09-04