Gemini 3.5 Flash-Lite tops out at nearly 350 tokens per second
OfficialLoganK · x · 2026-07-21
Google’s Gemini 3.5 Flash-Lite is being positioned as the company’s smallest and fastest Gemini model.
According to the post, it is:
- capable of nearly 350 output tokens per second
- often more intelligent than Gemini 3 in many cases
- priced the same as, but smarter than, Gemini 2.5 Flash
- faster than 3.1 Flash-Lite on most use cases
The poster says the latency makes it feel especially smooth for UI experiences and even viable for agent harnesses.
Related event: Google Launches Three New Gemini Models, Announces Gemini 4 Pre-training(157 posts)→
More from coding & agent
- Gergely Orosz: Shipping 10x PRs With AI Agents, Sites Fill With Small Regressions — ducha_aiki · 2026-09-11
- Same Echo Maze prompt, three frontier models: all passed visually but shipped the same hidden bug — eyishazyer · 2026-09-11
- Astra storyboards plus Minimax H3 per-shot generation boost video success rates — Hailuo_AI · 2026-09-11
- Codex tip: use Sol with Astra and Luna sub-agents to save usage — pvncher · 2026-09-11
- agents-best-practices: a provider-neutral Agent Skill for designing and auditing agentic harnesses — tom_doerr · 2026-09-11
- Cognition's SWE-2 uses a KKT duality argument in RL to shift the effort Pareto curve — YouJiacheng · 2026-09-11