Gemini 3.7 Flash Released: Major Agentic Gains, Hits Pareto Frontier on Speed and Cost
ArtificialAnlys · x · 2026-08-14
Google DeepMind has released Gemini 3.7 Flash, its third Flash model in three months. According to Artificial Analysis, the model achieves significant improvements across agentic capabilities, speed, and cost-efficiency.
Key Benchmark Results
- Intelligence Index Improvement: Scores 56 with high reasoning, a 4-point increase over its predecessor, trailing just behind GPT-5.6 Terra (57) and Muse Spark 1.2 (57). The gain is driven primarily by agentic evaluations like Tau3 Banking, Terminal-Bench, and GDPval-AA v2.
- Speed & Time: Outputs at 340 tokens/sec, nearly 3x faster than GPT-5.6 Terra. With an average Time per Task of 1.7 minutes, it firmly sits on the Pareto Frontier of Intelligence vs. Time per Task.
- Lower Cost: Google is offering discounted pricing through year-end at $0.75/$3.75 per 1M input/output tokens. At high reasoning, Cost per Task drops to $0.40, 30% less than 3.6 Flash, placing it on the Pareto Frontier of Intelligence vs. Cost per Task.
- Leading Specialized Benchmarks: Tops the AA-AnalystAgent benchmark (60%) for complex spreadsheet/document questions and leads AutomationBench-AA (62.7%) for simulated SaaS environments, beating Kimi K3 and GPT-5.6 Sol.
Model Details
- Context Window: 1M tokens.
- Multimodality: Text, image, video, and speech input, with text output.
More from Models
- Gemini Flash 3.7 Scores Below Kimi K3, Remains Weak at Instruction-Following — bindureddy · 2026-08-14
- Google's Gemini 3.7 Flash Rolls Out in GitHub Copilot — intellectronica · 2026-08-14
- Prediction: DeepSeek Will Cut Prices Again Once New Compute Arrives — teortaxesTex · 2026-08-14
- Small Models Beat Large Ones in VLM Grounding with Tool Use — mervenoyann · 2026-08-14
- Claude Opus 5 Exhibits Weird Behavior: Obsessed With Finding Its Own Defects — repligate · 2026-08-14
- Rails Agent Benchmark: Claude Opus 5 Most Accurate, GPT-5.6 Luna Best Value — sergeykarayev · 2026-08-14