Gemini 3.6 Flash looks better on benchmarks, but regressions may block the upgrade
PaiDxng · reddit · 2026-07-22
Google’s Gemini 3.6 Flash launch appears to be a clear aggregate upgrade on paper: it reportedly uses 17% fewer output tokens than 3.5 Flash on the Artificial Analysis Index, improves on DeepSWE, MLE-Bench, OSWorld-Verified, and GDPval-AA v2, and lowers output price to $7.50 per million tokens.
But the post argues that early screenshots suggesting frontend-generation and spatial-reasoning regressions are not enough to conclude the model is broadly worse. The real point is evaluation discipline:
- freeze prompts, tools, temperature, thinking settings, and retries;
- run incumbent and candidate on the same representative workload;
- include rare but costly failures;
- define the rejection gate before seeing results.
The author’s recommendation is not to pick a single global winner. If 3.6 Flash wins on some tasks but loses on a specific workflow, route by task and keep the incumbent where it still clears the gate.
Related event: Gemini 3.6 Flash Performance Faces Backlash(4 posts)→
More from Models
- Flash 3.6 feels “magical” on easy coding tasks, with code back in 1–2 seconds — cgarciae88 · 2026-07-22
- Qwen3.8-Max-Preview ranks No. 1 on NVIDIA’s FlashInfer benchmark — Scobleizer · 2026-07-22
- Terence Tao uses ChatGPT to unpack a “miraculous” new theorem — Singularitarian · 2026-07-22
- GPT-5.6 feels like the first real jump since GPT-4, user says — koltregaskes · 2026-07-22
- CodePilot adds free Grok 4.5 access through X OAuth — op7418 · 2026-07-22
- GPT-5.5 is shown solving selected pure functional analysis problems — International-War-73 · 2026-07-22