Gemini 3.6 Flash looks better on benchmarks, but regressions may block the upgrade

PaiDxng · reddit · 2026-07-22

Google’s Gemini 3.6 Flash launch appears to be a clear aggregate upgrade on paper: it reportedly uses 17% fewer output tokens than 3.5 Flash on the Artificial Analysis Index, improves on DeepSWE, MLE-Bench, OSWorld-Verified, and GDPval-AA v2, and lowers output price to $7.50 per million tokens.

But the post argues that early screenshots suggesting frontend-generation and spatial-reasoning regressions are not enough to conclude the model is broadly worse. The real point is evaluation discipline:

The author’s recommendation is not to pick a single global winner. If 3.6 Flash wins on some tasks but loses on a specific workflow, route by task and keep the incumbent where it still clears the gate.

Related event: Gemini 3.6 Flash Performance Faces Backlash(4 posts)→

Original post →

More from Models

Models channel →