Gemini 3.6 Flash Review: Faster and Cheaper, but Core Intelligence Unchanged

Google's newly released Gemini 3.6 Flash has sparked widespread community attention and hands-on testing. Official sources and multiple tests confirm a significant leap in efficiency: speed has increased by about twofold, pricing is down by 18%, and token consumption has been reduced. However, independent reviews generally note that its core intelligence level has not been upgraded accordingly, with uneven capabilities and even regressions in some areas, raising questions about its overall value for money.

Core Performance and Benchmarks

Regarding overall intelligence assessment, multiple authors point out that Gemini 3.6 Flash scored 50 on the Artificial Analysis (AA) index, exactly matching its predecessor, 3.5 Flash. @haider1 and @emax emphasize that while the new model shows improvements in specific leaderboards like coding (e.g., DeepSWE score rising from 37% to 49%), its overall performance still lags behind competitors like GPT-5.6 Sol and Terra. Furthermore, comparisons indicate it is 2.5 times more expensive than GPT-5.6 Luna. Additionally, tests by @skalskip92 and @llama_index reveal noticeable regressions in object detection and chart understanding (ParseBench dropped by 14%). @scaling01 also notes that in WeirdML tests, it often times out due to overly complex planning.

Strengths and Practical Experience

Despite no significant breakthrough in general intelligence, Gemini 3.6 Flash demonstrates robust capabilities in specific application scenarios. @allenainie mentions the model achieved a high score of 68% in the browser_use Web agent task test, surpassing GPT-5.6-sol and Sonnet 4.6. @mertdumenci highly praises its practical experience in search-oriented tasks, noting its response speed and actual performance far exceed the GPT or Claude series. @bytebot also confirms its multimodal capabilities remain solid, and web client usage is exceptionally fast.

Market Positioning and Controversy

Regarding the release strategy, @PaiDxng believes Google is attempting to drive user upgrades through better official benchmark numbers and token optimization, but regression testing results might hinder this progress. @burkov bluntly states that based on data from leaderboards like the Frontend Code Arena, Google has now fallen behind at least six labs in the top-tier model competition. Overall, Gemini 3.6 Flash is an iteration focused on efficiency and specific agent capabilities, rather than a leap in foundational intelligence.

2026-07-22 ~ 2026-07-23 · 16 related posts

Full story(4 episodes)→

Primary sources