Gemini 3.6 Flash Shows Stagnant Scores but Improved Efficiency

Artificial Analysis recently released full benchmark results for Google's new Gemini 3.6 Flash and 3.5 Flash-Lite. The data shows that while the new model brings no surprises in intelligence scores, it achieves significant optimizations in operational efficiency and cost control. With core performance stagnating, critics are beginning to question whether Google's R&D pace is falling behind in the LLM race.

Performance and Competitor Comparison

According to Artificial Analysis' leaderboard, Gemini 3.6 Flash scored 50 in intelligence with an Elo score of 1421. However, in key intelligence tests, its results are completely on par with the previous generation Gemini 3.5 Flash. Users @truecakesnake and @Rare_Bunch4348 pointed out that the model lacks substantial improvements and lags behind competitors like Meta Spark 1.1, GPT-5.6 Sol, and Grok 4.5 on the leaderboard. Additionally, @Angaisb_ noted that chart comparisons show Gemini 3.6 Flash is more expensive to use than GPT-5.6 Sol medium, yet delivers lower intelligence.

Efficiency Boosts and Cost Optimization

Despite no breakthroughs in absolute performance, Gemini 3.6 Flash shines in efficiency. According to data shared by @ArtificialAnlys, the model's output speed reaches approximately 304 token (note: the original post did not specify the exact time unit), and the average time per task has been halved compared to its predecessor. Meanwhile, the cost per task dropped from $0.59 to $0.50. Hands-on testing by @scaling01 corroborates this, suggesting that 3.6 Flash's Token efficiency is indeed slightly better than the 3.5 version.

Market Reception and Controversy

Faced with stagnant performance but improved efficiency, user opinions are divided. @Angaisb_ believes that whether the model is worth using depends entirely on practical efficiency; if efficiency isn't high enough, lowering the reasoning effort of GPT-5.6 Sol might be a better alternative. Meanwhile, @minxio_ and @iruletheworldmo compiled benchmark comparison tables, visually illustrating the comprehensive gap between Gemini 3.6 Flash and frontier models like GPT-5.6 Luna and Grok 4.5 across dimensions such as pricing, coding, and agentic tasks.

2026-07-21 ~ 2026-07-22 · 20 related posts

Full story(6 episodes)→