Gemini 3.7 Flash Benchmarks Leak: Major Gains but Instruction Following Lags
Recent leaks reveal comprehensive benchmark scores for Google's Gemini 3.7 Flash. The data shows massive improvements over its predecessor in coding, reasoning, and text arenas, redefining cost-efficiency with blazing speed and rock-bottom pricing. However, reviewers note that instruction following remains a weak point, and an "overthinking" issue could unexpectedly drive up costs.
Confirmed
- Arena Scores: According to LMSYS data, Gemini 3.7 Flash (High) scored 1490 on the Text Arena, ranking 9th—a significant jump from 3.6 Flash's 16th place. It also topped the Code Arena: WebDev with a score of 1588.
- Benchmark Improvements: OSWorld score jumped from 33.8% to 47.9%; FrontierCode benchmark increased from 34.4% to 43.6%; knowledge work reasoning (GDP.pdf benchmark) rose from 22% to 34%.
- Intelligence and Cost-Efficiency: It scored 56 on the Artificial Analysis Intelligence Index (category median is 34) and topped recommendations requiring both "fastest speed" and "highest intelligence." Creator @petrusenkomax confirmed its cost is halved compared to the previous generation.
- Suboptimal Areas: Ranked 7th in the VoxelBench evaluation; developer Zachary Nado pointed out it fell short of SOTA levels on the BenchBench evaluation.
Unconfirmed
- Real-World Experience Controversies: Reviewer @bindureddy noted its score was slightly below Kimi K3, and weak instruction following led to a subpar practical experience. Furthermore, @teortaxesTex criticized the model for "overthinking," making it underperform compared to the Pro version in certain use cases while costing up to 4.5 times more, resulting in poor cost-efficiency.
Why It Matters
Gemini 3.7 Flash demonstrates Google's aggressive pricing and performance leaps in the mid-range lightweight model market, directly raising the cost-efficiency bar for competitors. Yet, the disconnect between benchmark scores and real-world instruction-following capabilities has reignited debates among developers regarding the validity of current evaluation systems.
2026-08-13 ~ 2026-08-14 · 14 related posts
- Episode 1: Gemini 3.7 Flash Benchmarks Leak: Major Gains but Instruction Following Lags(2026-08-13, 14 posts)
- Episode 2: Gemini 3.7 Flash Surfaces with Halved Pricing as Pro Version Rumored to Be Shelved(2026-08-13, 9 posts)
- Episode 3: Google Launches Gemini 3.7 Flash: Major Coding Upgrades and 50% Price Cut(2026-08-14, 40 posts)
Primary sources
- Gemini Flash Slammed for Poor Cost-Performance: 4.5x Costlier Than Pro in Trials — teortaxesTex · 2026-08-13
- Gemini 3.7 Flash Benchmark Results Leaked Online — virtualQubit · 2026-08-14
- Gemini 3.7 Flash Benchmark Results Revealed — Expensive_Syrup_6529 · 2026-08-14
- Gemini 3.7 Flash Jumps to #9 on Text Arena with Major Category Gains — arena · 2026-08-14
- Gemini 3.7 Flash Ranks 7th on VoxelBench Amid Pro Model Anticipation — legit_api · 2026-08-14
- [source] Gemini 3.7 Flash Scores 1490 Points on the Text Arena — arena · 2026-08-14
- Gemini 3.7 Flash Tops Code Arena, Reshaping Pareto Frontier with $0.75 Pricing — airesearch12 · 2026-08-14
- [source] Gemini 3.7 Flash Benchmarks: Ranked #1 in Speed and #16 in Intelligence — ArtificialAnlys · 2026-08-14
- Gemini 3.7 Flash Review: Strong Performance but Not SOTA — zacharynado · 2026-08-14
- Gemini 3.7 Flash Tested: Major Boosts in Coding and Reasoning at Half the Cost — petrusenko_max · 2026-08-14
- Leaked Gemini 3.7 Flash Benchmarks Show Major Coding Gains, Beating Sonnet 5 at Lower Cost — ChrisGPT · 2026-08-14
- Debunking 'Google is Cooked': Gemini Flash Tops Speed and Intelligence Charts — altryne · 2026-08-14
- [source] Gemini Flash 3.7 Scores Below Kimi K3, Remains Weak at Instruction-Following — bindureddy · 2026-08-14
- Gemini Flash 3.7 scores just below Kimi K3, but instruction-following remains weak — bindureddy · 2026-08-14