Building live voice agents with Gemini: DeepMind demos Live API, Lyria 3, real-time translation
AI Engineer · youtube · 2026-10-11
Thor Schaeff of Google DeepMind developer experience demos how to build agents that hear, see, and answer in your language in real time using Gemini.
Highlights
- An illustrated story built with Gemini Live and the fast Nano Banana image model, using the stateful Interactions API to keep characters and style consistent across images
- Near real-time speech translation with Gemini 3.5 Live Translate, letting the audience hear the talk in their own language via QR code
- An AI radio DJ voice agent that calls tools, using Lyria 3 to generate a reggaeton track on request
- How the Live API works: a stateful WebSocket session taking text, audio, and video, streaming back audio plus transcript — a native audio model rather than an STT+TTS pipeline, with 90 languages, barge-in, tool calls, and Google Search grounding
- Closing voice + vision demo (including Mandarin) in Google AI Studio; partners LiveKit and Pipecat
Docs linked for the Live API, Interactions API, Live Translate, and Lyria 3.
More from coding & agent
- Running GLM-5.3-Flash on dual Ascend 310P cards: 8-9 tok/s and 311K context — matteiuspi · 2026-10-12
- plain writing is the team's single most-used internal skill, by a factor of 2 — sh_reya · 2026-10-12
- plain-writing-skill: an open-source skill that makes AI agents write plainly — sh_reya · 2026-10-12
- User says Grokbot autonomously won new business, calling Opus + harness "magical" — iruletheworldmo · 2026-10-12
- Jose Valim shows Campfire AI benchmark was rigged: Elixir faced far stricter checks than Go/Rust — zeeg · 2026-10-12
- Ask the model for bullet points, write the changes yourself: keeping your voice with AI — ctjlewis · 2026-10-12