Gemini Batch API latency drops 80% as Google adds partial batch support
OfficialLoganK · x · 2026-07-21
Google says it shipped major infrastructure upgrades for the Gemini Batch API:
- p95 latency is down 80%
- p99 latency is down 68%
- batch success rate is now above 99.998%
- batch expirations are down 98%
- partial batch support has been added
The post frames this as a team win on reliability and throughput for batch workloads.
Related event: Google Upgrades Gemini Batch API with 80% Latency Drop(2 posts)→
More from Infra
- SF Compute founder: buying compute is 'an absolutely awful experience' right now — IgorCarron · 2026-09-11
- SmolVM open-sources persistent computer infrastructure for agents that outlive chat sessions — aniketmaurya · 2026-09-11
- PyTorch Day Korea 2026 launches first offline conf, CFP closes Sept 13 — PyTorch · 2026-09-11
- Local LLM server dilemma: 4x CMP-170HX (price up 53% in 20 days) vs Mac Studio M5 Ultra — rumboll · 2026-09-11
- llama.cpp lands Flash Attention tuning for RDNA4, big prefill gains on AMD — pmttyji · 2026-09-11
- Your p99 latency benchmark may be lying: a deep dive into coordinated omission — Franc0Fernand0 · 2026-09-11