Gemini 3.8 Flash outputs ~300 tokens/s, needs 2.5 min per task on high reasoning
ArtificialAnlys · x · 2026-09-02
Artificial Analysis published speed and efficiency data for Gemini 3.8 Flash:
- On high reasoning it averages 300 output tokens/second with a Time per Task of 2.5 minutes, just ahead of GPT-5.6 Luna (max, 2.6 min) and GPT-5.6 Terra (max, 3.3 min)
- Higher token usage vs 3.7 Flash pushes Time per Task from 2.2 to 2.5 minutes, behind Claude Fable 5.1 (medium, 2.1 min)
- On low reasoning it drops to 0.8 minutes per task, landing on the Intelligence vs. Time per Task Pareto frontier
Related event: Gemini 3.8 Flash Benchmarks: Smarter, Pricier, Still on the Cost Frontier(8 posts)→
More from Models
- GPT-6-Astra model slug spotted on OpenAI APIs, hinting at routing tests — testingcatalog · 2026-09-03
- Polymarket puts 78% odds on OpenAI's rumored Astra model launching tomorrow — Polymarket · 2026-09-03
- Polymarket puts 78% odds on OpenAI's rumored Astra model shipping tomorrow — Polymarket · 2026-09-03
- Unverified report: Gemini 3.8 Flash out with big agentic coding gains — clmt · 2026-09-03
- 'GPT-6-ASTRA' spotted staged on the OpenAI API, unconfirmed — ThunderBeanage · 2026-09-03
- Google DeepMind releases Gemini 3.8 Flash and 3.8 Flash Cyber — Google DeepMind · 2026-09-03