Google launches Gemini 3.6 Flash with lower token use, plus a faster Flash-Lite tier
aniketmaurya · x · 2026-07-22
Google has rolled out Gemini 3.6 Flash and Gemini 3.5 Flash-Lite.
- Gemini 3.6 Flash is positioned as the efficient workhorse for coding, knowledge work, and agentic workflows. Google says it uses 17% fewer output tokens than Gemini 3.5 Flash, with up to 65% token savings on some benchmarks. Pricing is $1.50 / 1M input tokens and $7.50 / 1M output tokens. Reported gains include 83.0% on OSWorld Verified and 49% on DeepSWE.
- Gemini 3.5 Flash-Lite targets high-throughput, low-latency agent tasks and is said to reach up to 350 output tokens/sec. Pricing is $0.30 / 1M input tokens and $2.50 / 1M output tokens.
Google says both models are rolling into the Gemini app, with API access in Google AI Studio and Android Studio. A limited-access Gemini 3.5 Flash Cyber pilot is also coming via CodeMender.
Related event: Google Launches Gemini 3.6 Flash and Other New Models(59 posts)→
More from coding & agent
- Loop vs Graph Engineering: The Architectural Shift in AI Agents — ghumare64 · 2026-07-22
- FactoryAI gave back its first millions, then shipped Droid CLI two years later — matanSF · 2026-07-22
- Hermes Agent Refactoring Proposal: Decoupling via Event Bus and Monorepo Slicing — Promptmethus · 2026-07-22
- Show HN: a 6.2 MB pure-Go terminal command palette with no fzf dependency — MarinhoD · 2026-07-22
- ty now reads Pydantic config keywords and field metadata — charliermarsh · 2026-07-22
- Pensar Launches AI Security Agent to Autonomously Discover and Patch 0-Days — andriy_mulyar · 2026-07-22