Gemini 3.6 Flash Review: Faster and Cheaper, But Not Smarter

Google's latest release, Gemini 3.6 Flash, has sparked widespread testing and discussion within the community. The model achieves significant efficiency improvements, operating about twice as fast, reducing prices by 18%, and consuming fewer tokens. However, independent reviews generally note that its core intelligence has not been upgraded; its overall performance remains on par with the previous generation, with regressions in some capabilities, raising questions about its cost-effectiveness and upgrade value.

Confirmed

In terms of overall intelligence assessment, multiple authors point out that Gemini 3.6 Flash scores 50 on the Artificial Analysis (AA) Index, exactly matching its predecessor. @haider1 and @emax emphasize that although the new model shows progress in coding benchmarks (e.g., DeepSWE score increased from 37% to 49%), its overall performance still lags behind competitors like GPT-5.6 Sol and Terra, and it is criticized for being 2.5 times more expensive than GPT-5.6 Luna. Tests by @skalskip92 and @llamaindex reveal noticeable regressions in object detection and chart understanding (ParseBench dropped 14%), and @scaling01 notes that it often times out in WeirdML tests due to overly complex planning. In specific application scenarios, the model demonstrates strong capabilities: @allenainie mentions it achieved a high score of 68% in the browseruse web agent task test, surpassing GPT-5.6-sol and Sonnet 4.6; @mertdumenci highly praises its implementation experience in search-oriented tasks; and @bytebot confirms its multimodal capabilities remain solid and its web interface is extremely fast.

Unconfirmed

There is divided community feedback on whether the model feels smarter in subjective experience. @sankineth felt the model was "significantly smarter" after a brief trial, but this individual subjective experience differs from the "intelligence plateau" conclusion drawn from most benchmark tests.

Why it matters

The release strategy of this model reflects an adjustment in Google's roadmap. @PaiDxng believes Google is attempting to drive user upgrades through better official benchmark numbers and token optimization, but regression test results might hinder this process. A third-party report shared by @sujingshen indicates that Google is shifting its strategy towards reducing costs by 30-40% while maintaining capability parity. However, @burkov bluntly states that, based on data from arenas like the Frontend Code Arena, Google has now been surpassed by at least 6 other labs in the top-tier model competition. Overall, Gemini 3.6 Flash is an iteration focused on efficiency and specific agent capabilities rather than a leap in foundational intelligence.

2026-07-22 ~ 2026-07-23 · 16 related posts

Full story(5 episodes)→

Primary sources