Gemini 3.8 Flash: same token price, different cost per task — a real-world billing test
JuggernautCritical92 · reddit · 2026-09-04
Google says Gemini 3.8 Flash keeps the same input/output token prices as 3.7 Flash, but 3.8 may take extra reasoning steps and call tools more often on hard tasks—so a flat token price can still mean a different bill per completed task.
Google's advice: lower the effort setting or stay on 3.7 Flash when efficiency matters; planning and difficult code changes may benefit from the extra work, while classifiers and routine extraction may just burn more tokens.
The author tested both slugs through the same ZenMux API setup with identical tasks and timeouts, tracking retries and accepted patches. Both did the work well with no meaningful difference—but the author admits the tasks were too easy. The real question is whether 3.8 completes enough extra tasks to cover its extra reasoning, which requires accepted results and total tokens from the same runs, not a launch-page price.
More from Infra
- Morgan Stanley: 20% of Star Market IPOs now target China's 'chokepoint' tech — pstAsiatech · 2026-09-04
- DeepSeek plans 160,000-chip Huawei Ascend 950DT cluster in Inner Mongolia, Bloomberg reports — kimmonismus · 2026-09-04
- Zeiss executive: China about 15 years behind on EUV lithography tools — pstAsiatech · 2026-09-04
- Minisforum MS-S1 MAX P495 listing hints at €7999 price, double the original MS-S1 MAX — fairydreaming · 2026-09-04
- First community MLX 4-bit benchmarks for K2-Horizon-MoVA-36B hit 49.1 tok/s locally — DerTomsn · 2026-09-04
- TSMC doubles equipment forecast in six months, builds 20 fabs to chase AI demand — firstadopter · 2026-09-04