Google’s Gemini 3.6 Flash trims token use as 3.5 Flash-Lite targets low cost
赛博禅心 · wechat · 2026-07-21
Gemini 3.6 Flash cuts token use, while 3.5 Flash-Lite targets ultra-low cost
The post says Google has just released three models:
- Gemini 3.6 Flash: an upgrade to 3.5 Flash, with token usage reduced by 17% overall and by up to 65% on DeepSWE; priced at $1.50 input / $7.50 output per 1M tokens.
- Gemini 3.5 Flash-Lite: positioned as the fastest and cheapest model in the 3.5 family, reaching 350 output tokens/s; priced at $0.30 input / $2.50 output per 1M tokens.
- Gemini 3.5 Flash Cyber: a cybersecurity-specialized model fine-tuned from 3.5 Flash and limited to governments and trusted partners.
The post also claims Gemini 3.5 Pro is in internal testing and Gemini 4 has started training.
Related event: Google Launches Multiple Gemini Models for Enhanced Cost-Efficiency(67 posts)→
More from Models
- NVIDIA says Nemotron 3 Ultra scored 30/42 on the 2026 IMO problems — NVIDIAAI · 2026-07-22
- Gemma-4-26B-a4B reportedly beats Qwen3.6 and Qwen3.5 MoE fine-tunes — JLeonsarmiento · 2026-07-22
- OpenAI is reportedly briefing U.S. lawmakers on its next model family — kimmonismus · 2026-07-22
- Muse Spark 1.1 lands at 1495 on Text Arena with standout agentic-coding price performance — ycombinator · 2026-07-22
- Advanced AI Models Are Becoming Impossible to Plug and Play — emollick · 2026-07-22
- Google Gemini's AI Problem: No Leading Model for Core Workloads — bindureddy · 2026-07-22