Google's Gemini 3.8 Flash 'works harder' but may burn more tokens at same pricing
The Verge AI · rss · 2026-09-03
Just weeks after Gemini 3.7 Flash, Google launched Gemini 3.8 Flash, claiming it "works harder" by running more reasoning steps and calling tools iteratively on complex tasks. Intro pricing matches 3.7 Flash: $0.75/M input and $3.75/M output tokens. But Google warns the model may use more tokens to maximize performance, especially at higher effort levels, so costs could climb. Developers wanting to minimize token usage can stay on 3.7 Flash. Early reactions to the launch have been mixed.
More from Infra
- Perplexity Computer demos fully local operation on NVIDIA DGX Spark — chrmanning · 2026-09-03
- llama.cpp deprecates --chat-template-kwargs, reasoning-preserve now on by default — Bulky-Priority6824 · 2026-09-03
- Agentic API adds a stateful layer in front of vLLM for open-model agent runtimes — techNmak · 2026-09-03
- Mitchell Hashimoto Details Memory Optimization Tricks in the Superlogical Server — sull · 2026-09-03
- Perplexity's Lily beats MLX-LM with 1.23x prefill and 1.35x decode throughput on M5 Max — perplexity_ai · 2026-09-03
- Perplexity open-sources Lily, a local inference engine for Qwen3.6 on Apple silicon — perplexity_ai · 2026-09-03