Gemini 3.6 Flash looks better on benchmarks, but regressions may block the upgrade
PaiDxng · reddit · 2026-07-22
Google’s Gemini 3.6 Flash launch appears to be a clear aggregate upgrade on paper: it reportedly uses 17% fewer output tokens than 3.5 Flash on the Artificial Analysis Index, improves on DeepSWE, MLE-Bench, OSWorld-Verified, and GDPval-AA v2, and lowers output price to $7.50 per million tokens.
But the post argues that early screenshots suggesting frontend-generation and spatial-reasoning regressions are not enough to conclude the model is broadly worse. The real point is evaluation discipline:
- freeze prompts, tools, temperature, thinking settings, and retries;
- run incumbent and candidate on the same representative workload;
- include rare but costly failures;
- define the rejection gate before seeing results.
The author’s recommendation is not to pick a single global winner. If 3.6 Flash wins on some tasks but loses on a specific workflow, route by task and keep the incumbent where it still clears the gate.
Related event: Gemini 3.6 Flash Review: Faster and Cheaper, But Not Smarter(16 posts)→
More from Models
- Meta's Muse Agent has built-in invite code logic, hinting at free-usage expansion — testingcatalog · 2026-09-11
- Same Echo Maze prompt, three frontier models: all passed visually but shipped the same hidden bug — eyishazyer · 2026-09-11
- Benchmark scores drop from 89% to 19% on new evals — how benchmaxxing breaks leaderboard trust — airesearch12 · 2026-09-11
- ChatGPT tells user their question is too hard and to 'accept dumber answers' — phido3000 · 2026-09-11
- Claude is no longer available for minors as Anthropic rolls out age assurance — Muhammad523 · 2026-09-11
- Developer Building a Unified Leaderboard of All Model Benchmark Scores — airesearch12 · 2026-09-11