Gemini 3.6 Flash Shows Stagnant Scores but Improved Efficiency
Artificial Analysis recently released full benchmark results for Google's new Gemini 3.6 Flash and 3.5 Flash-Lite. The data shows that while the new model brings no surprises in intelligence scores, it achieves significant optimizations in operational efficiency and cost control. With core performance stagnating, critics are beginning to question whether Google's R&D pace is falling behind in the LLM race.
Performance and Competitor Comparison
According to Artificial Analysis' leaderboard, Gemini 3.6 Flash scored 50 in intelligence with an Elo score of 1421. However, in key intelligence tests, its results are completely on par with the previous generation Gemini 3.5 Flash. Users @truecakesnake and @Rare_Bunch4348 pointed out that the model lacks substantial improvements and lags behind competitors like Meta Spark 1.1, GPT-5.6 Sol, and Grok 4.5 on the leaderboard. Additionally, @Angaisb_ noted that chart comparisons show Gemini 3.6 Flash is more expensive to use than GPT-5.6 Sol medium, yet delivers lower intelligence.
Efficiency Boosts and Cost Optimization
Despite no breakthroughs in absolute performance, Gemini 3.6 Flash shines in efficiency. According to data shared by @ArtificialAnlys, the model's output speed reaches approximately 304 token (note: the original post did not specify the exact time unit), and the average time per task has been halved compared to its predecessor. Meanwhile, the cost per task dropped from $0.59 to $0.50. Hands-on testing by @scaling01 corroborates this, suggesting that 3.6 Flash's Token efficiency is indeed slightly better than the 3.5 version.
Market Reception and Controversy
Faced with stagnant performance but improved efficiency, user opinions are divided. @Angaisb_ believes that whether the model is worth using depends entirely on practical efficiency; if efficiency isn't high enough, lowering the reasoning effort of GPT-5.6 Sol might be a better alternative. Meanwhile, @minxio_ and @iruletheworldmo compiled benchmark comparison tables, visually illustrating the comprehensive gap between Gemini 3.6 Flash and frontier models like GPT-5.6 Luna and Grok 4.5 across dimensions such as pricing, coding, and agentic tasks.
2026-07-21 ~ 2026-07-22 · 20 related posts
- Episode 1: Rumor: Google to Launch Gemini 3.6 Flash with Lower Price and Higher Scores(2026-07-21, 8 posts)
- Episode 2: Gemini 3.6 Flash Glitch: Misidentifies Google's Latest Model(2026-07-21, 2 posts)
- Episode 3: Google Launches Gemini 3.6 Flash and Other New Models(2026-07-21, 64 posts)
- Episode 4: Gemini 3.6 Flash Shows Stagnant Scores but Improved Efficiency(2026-07-21, 20 posts)
- Episode 5: Google Starts Gemini 4 Pre-training; 3.5 Pro in Partner Testing(2026-07-21, 13 posts)
- Episode 6: Google Launches Gemini 3.5 Flash Cyber Security Model(2026-07-22, 2 posts)
- Gemini 3.6 Flash will only matter if it is extremely efficient — Angaisb_ · 2026-07-21
- Reddit gallery compares Gemini 3.5 Flash-Lite and 3.6 Flash benchmarks — NaM_777 · 2026-07-21
- Artificial Analysis ranks Gemini 3.6 Flash at 50 on its updated intelligence index — Angaisb_ · 2026-07-21
- Gemini 3.6 Flash is pricier than GPT-5.6 Sol medium, chart claims — Angaisb_ · 2026-07-21
- Benchmark: Gemini 3.6 Flash Scores Parity with 3.5 Flash — truecakesnake · 2026-07-21
- Artificial Analysis chart compares Gemini 3.5 Flash-Lite with 3.6 Flash — Expensive_Syrup_6529 · 2026-07-21
- Hands-on: Gemini 3.6 Flash is Slightly More Token-Efficient Than 3.5 — scaling01 · 2026-07-21
- Gemini 3.6 Flash matches 3.5 Flash on Artificial Analysis and trails newer rivals — Rare_Bunch4348 · 2026-07-21
- Gemini 3.6 Flash matches 3.5 Flash on the same intelligence score — Hesamation · 2026-07-21
- [source] Gemini 3.6 Flash halves task time while 3.5 Flash-Lite gets faster but pricier — ArtificialAnlys · 2026-07-21
- [source] Gemini 3.6 Flash gets cheaper per task, while 3.5 Flash-Lite more than doubles in cost — ArtificialAnlys · 2026-07-21
- Artificial Analysis puts Gemini 3.6 Flash at 1421 Elo on GDPval-AA v2 — ArtificialAnlys · 2026-07-21
- Full benchmark results surface for Gemini 3.6 Flash and 3.5 Flash-Lite — ArtificialAnlys · 2026-07-21
- [source] Artificial Analysis shows Gemini 3.6 Flash and 3.5 Flash-Lite improve on agentic work — ArtificialAnlys · 2026-07-21
- Benchmark scorecard pits GPT-5.6 Sol, Claude Fable 5, and Gemini 3.6 Flash — iruletheworldmo · 2026-07-22
- Gemini 3.6 Flash matches 3.5 Flash on Artificial Analysis benchmark — airesearch12 · 2026-07-22
- Gemini 3.6 Flash benchmark results reignite concerns that Google is slipping behind — minxio_ · 2026-07-22
- Benchmark chart pits GPT-5.6 Luna, Grok 4.5 and Gemini 3.6 Flash on price and scores — iruletheworldmo · 2026-07-22
- Gemini 3.6 Flash ties 3.5 Flash on intelligence while cutting task time — ziv_ravid · 2026-07-22
- Gemini 3.6 Flash looks pricier than Grok 4.5 on the same task — XFreeze · 2026-07-22