FULL STORY

Gemini 3.7 Flash: From Backend Leak to Launch and Benchmarks

Following backend leaks and rumors of halved pricing, Google officially launched Gemini 3.7 Flash. The model delivers major improvements in coding and agentic tasks, with benchmarks confirming its strong cost-effectiveness.

2026-08-13 ~ 2026-08-14 · 3 episodes · 48 posts

Episode 1 · Gemini 3.7 Flash Surfaces with Halved Pricing as Pro Version Rumored to Be Shelved (2026-08-13, 9 posts)

Multiple leaks indicate that Google's Gemini 3.7 Flash has appeared in the Google Cloud Console and Enterprise backend, with an expected release today. Pricing has reportedly been significantly reduced to $0.75/million tokens for input and $3.75/million tokens for output. Meanwhile, rumors suggest that 3.5 Pro has been internally deprecated, and the Pro version might wait until 4.0.

已确认

  • According to @RareBunch4348 and @scaling01 (reposting), Gemini 3.7 Flash has appeared in the Google Cloud Console,暗示发布在即。
  • @testingcatalog leaked that Gemini 3.7 Flash has surfaced in the Gemini Enterprise backend with the model ID gemini-3.7-flash, currently in a hidden state.
  • @adonissingh leaked that Gemini 3.7 Flash is expected to be released later today, with pricing drastically lowered to $0.75/million tokens for input and $3.75/million tokens for output.

尚未确认

  • @adonissingh stated that Gemini 3.5 Pro has been internally deprecated but provided no further details.
  • @koltregaskes speculated that the Pro version might not launch until Gemini 4.0, which remains a guess.

为什么重要

The release of Gemini 3.7 Flash will introduce lower pricing, potentially impacting AI model market competition. At the same time, the deprecation of 3.5 Pro and the delay of the Pro version to 4.0 could affect user expectations regarding the pace of Google's model iterations.

Episode 2 · Google Launches Gemini 3.7 Flash with Upgraded Coding and Half-Price Cost (2026-08-14, 31 posts)

Google officially launched the Gemini 3.7 Flash model, marking a major leap in coding and agentic workflows. The model is available at a 50% discount for API calls before year-end. This signifies a substantial breakthrough in the practicality of large models for complex automation and software development, making it a key focus for developers.

已确认

  • 核心升级: Gemini 3.7 Flash is specifically optimized for coding and agentic tasks. In key benchmarks, the DeepSWE score surged from 49.0% to 65.3%, and FrontierCode also saw improvements from 34.4%.
  • 定价与可用性: API call prices are 50% cheaper than the previous 3.6 Flash generation until the end of the year. The model is fully rolled out across the API, AI Studio, and Antigravity platforms, and is integrated into consumer-facing Google AI Pro and Ultra subscriptions.
  • 智能体表现: Developer @philschmid noted in tests that the new model exhibits more discipline in agentic workflows. It explores, parses errors, and runs tests before modifying code, effectively reducing redundant turns.
  • 硬核演示: To showcase its coding accuracy and agentic capabilities, the Google team demonstrated building 3 agent teams to fully autonomously train a robot control model from scratch.

为什么重要

  • 迭代速度惊人: According to @koraykv and @cyruszei, the model iterated from 3.5 to 3.7 in just 3 months (only three weeks after 3.6's release). Driven by developer feedback and algorithmic innovation, this demonstrates Google's extremely rapid pace of technological advancement.
  • 性价比重塑: By halving the price while significantly improving complex task processing, it directly lowers the compute cost barrier for enterprise-level development and agent construction.

11 more related posts →

Episode 3 · Gemini 3.7 Flash Benchmarks Leak: Arena Surge and Pareto Frontier Shift (2026-08-14, 8 posts)

Recent leaks have revealed a slew of benchmark and arena scores for Google's Gemini 3.7 Flash (High) model. The data indicates significant capability leaps, redefining the Pareto frontier of large language model performance and pricing with its outstanding cost-effectiveness, which has sparked widespread community interest.

Confirmed

  • Arena Score Surge: According to LMSYS Chatbot Arena data, Gemini 3.7 Flash (High) scored 1490 points in the text arena, ranking 9th—a substantial jump from its predecessor 3.6 Flash (High)'s 16th place. Poster @arena noted marked improvements across all sub-categories.
  • Standout Coding Capabilities: The model achieved an impressive 1588 score in Code Arena: WebDev, praised for redefining the price-to-performance Pareto frontier.
  • Other Benchmarks: It secured 7th place in the VoxelBench evaluation, a result acknowledged by poster @legitapi.

Why It Matters

  • Cost-Performance Benchmark: Maintaining its lightweight positioning, Gemini 3.7 Flash demonstrates arena prowess that rivals or even surpasses some flagship models, offering developers a highly cost-effective new option.
  • Heightened Expectations: The robust competitiveness of this lightweight Flash version has significantly raised the community's anticipation for Google's upcoming Pro-tier models.