LiveKit says Gemma 4 31B hits 192 ms to first token in voice agents
GlennCameronjr · x · 2026-07-30
LiveKit says the fastest LLM for a voice agent right now is an open-weight model: Gemma 4 31B on LiveKit Inference reaches first token in 192 ms and a full spoken sentence in 354 ms, while costing $0.40 per 1M input tokens ($0.20 cached) and $1.20 per 1M output tokens.
The post frames this as a practical latency-plus-cost win for real-time voice agents, not just a model-quality comparison.
More from Infra
- Replacing Cloud Vision APIs Locally with Nvidia Nemotron on DGX Spark — JFPuget · 2026-07-30
- Laguna XS Hits Apple Silicon, Doubling Decode Speed to Over 140 TPS — gajesh · 2026-07-30
- Laguna XS Breaks 140 TPS on Apple Machines with New FAST Mode — gajesh · 2026-07-30
- $50B+ in AI Data Center Leases Signed in July as Bitcoin Miners Pivot — abhiadesai · 2026-07-30
- 26B-Parameter Gemma 4 Runs on Mac with 2GB RAM via SSD Streaming — petrusenko_max · 2026-07-30
- Cloudflare Containers Introduces .exec(): Stream Request Body Directly to Container Process — craigsdennis · 2026-07-30