FULL STORY

Gemini 3.7 Flash: From Leaked Benchmarks to Launch

Google's Gemini 3.7 Flash transitioned from leaked benchmarks to official launch on August 14. The newly released model offers a 50% price cut and significant performance upgrades for coding workflows.

2026-08-13 ~ 2026-08-14 · 3 episodes · 65 posts

Episode 1 · Gemini 3.7 Flash Benchmarks Leak: Big Gains, but Instruction Following Still Weak (2026-08-13, 15 posts)

Recent leaks reveal comprehensive benchmark scores for Google's Gemini 3.7 Flash. The data shows massive improvements over its predecessor in coding, reasoning, and text arenas, redefining cost-efficiency with blazing speed and rock-bottom pricing. However, reviewers note that instruction following remains a weak point, and an "overthinking" issue could unexpectedly drive up costs.

Confirmed

  • Arena Scores: According to LMSYS data, Gemini 3.7 Flash (High) scored 1490 on the Text Arena, ranking 9th—a significant jump from 3.6 Flash's 16th place. It also topped the Code Arena: WebDev with a score of 1588.
  • Benchmark Improvements: OSWorld score jumped from 33.8% to 47.9%; FrontierCode benchmark increased from 34.4% to 43.6%; knowledge work reasoning (GDP.pdf benchmark) rose from 22% to 34%.
  • Intelligence and Cost-Efficiency: It scored 56 on the Artificial Analysis Intelligence Index (category median is 34) and topped recommendations requiring both "fastest speed" and "highest intelligence." Creator @petrusenkomax confirmed its cost is halved compared to the previous generation.
  • Suboptimal Areas: Ranked 7th in the VoxelBench evaluation; developer Zachary Nado pointed out it fell short of SOTA levels on the BenchBench evaluation.

Unconfirmed

  • Real-World Experience Controversies: Reviewer @bindureddy noted its score was slightly below Kimi K3, and weak instruction following led to a subpar practical experience. Furthermore, @teortaxesTex criticized the model for "overthinking," making it underperform compared to the Pro version in certain use cases while costing up to 4.5 times more, resulting in poor cost-efficiency.

Why It Matters

Gemini 3.7 Flash demonstrates Google's aggressive pricing and performance leaps in the mid-range lightweight model market, directly raising the cost-efficiency bar for competitors. Yet, the disconnect between benchmark scores and real-world instruction-following capabilities has reignited debates among developers regarding the validity of current evaluation systems.

Episode 2 · Gemini 3.7 Flash Surfaces with Halved Pricing as Pro Version Rumored to Be Shelved (2026-08-13, 9 posts)

Multiple leaks indicate that Google's Gemini 3.7 Flash has appeared in the Google Cloud Console and Enterprise backend, with an expected release today. Pricing has reportedly been significantly reduced to $0.75/million tokens for input and $3.75/million tokens for output. Meanwhile, rumors suggest that 3.5 Pro has been internally deprecated, and the Pro version might wait until 4.0.

已确认

  • According to @RareBunch4348 and @scaling01 (reposting), Gemini 3.7 Flash has appeared in the Google Cloud Console,暗示发布在即。
  • @testingcatalog leaked that Gemini 3.7 Flash has surfaced in the Gemini Enterprise backend with the model ID gemini-3.7-flash, currently in a hidden state.
  • @adonissingh leaked that Gemini 3.7 Flash is expected to be released later today, with pricing drastically lowered to $0.75/million tokens for input and $3.75/million tokens for output.

尚未确认

  • @adonissingh stated that Gemini 3.5 Pro has been internally deprecated but provided no further details.
  • @koltregaskes speculated that the Pro version might not launch until Gemini 4.0, which remains a guess.

为什么重要

The release of Gemini 3.7 Flash will introduce lower pricing, potentially impacting AI model market competition. At the same time, the deprecation of 3.5 Pro and the delay of the Pro version to 4.0 could affect user expectations regarding the pace of Google's model iterations.

Episode 3 · Google Launches Gemini 3.7 Flash with 50% Price Cut (2026-08-14, 41 posts)

Google officially launched the Gemini 3.7 Flash model, marking a major leap in coding and agentic workflows. The model is available at a 50% discount for API calls before year-end. This signifies a substantial breakthrough in the practicality of large models for complex automation and software development, making it a key focus for developers.

已确认

  • 核心升级: Gemini 3.7 Flash is specifically optimized for coding and agentic tasks. In key benchmarks, the DeepSWE score surged from 49.0% to 65.3%, and FrontierCode also saw improvements from 34.4%.
  • 定价与可用性: API call prices are 50% cheaper than the previous 3.6 Flash generation until the end of the year. The model is fully rolled out across the API, AI Studio, and Antigravity platforms, and is integrated into consumer-facing Google AI Pro and Ultra subscriptions.
  • 智能体表现: Developer @philschmid noted in tests that the new model exhibits more discipline in agentic workflows. It explores, parses errors, and runs tests before modifying code, effectively reducing redundant turns.
  • 硬核演示: To showcase its coding accuracy and agentic capabilities, the Google team demonstrated building 3 agent teams to fully autonomously train a robot control model from scratch.

为什么重要

  • 迭代速度惊人: According to @koraykv and @cyruszei, the model iterated from 3.5 to 3.7 in just 3 months (only three weeks after 3.6's release). Driven by developer feedback and algorithmic innovation, this demonstrates Google's extremely rapid pace of technological advancement.
  • 性价比重塑: By halving the price while significantly improving complex task processing, it directly lowers the compute cost barrier for enterprise-level development and agent construction.

21 more related posts →