LiveKit says Gemma 4 31B hits 192 ms to first token in voice agents

GlennCameronjr · x · 2026-07-30

LiveKit says the fastest LLM for a voice agent right now is an open-weight model: Gemma 4 31B on LiveKit Inference reaches first token in 192 ms and a full spoken sentence in 354 ms, while costing $0.40 per 1M input tokens ($0.20 cached) and $1.20 per 1M output tokens.

The post frames this as a practical latency-plus-cost win for real-time voice agents, not just a model-quality comparison.

Original post →

More from Infra

Infra channel →