Google ships Gemini 3.8 Live speech-to-speech models; Willison builds a zero-dependency web UI
Simon Willison · rss · 2026-09-16
Google released Gemini 3.8 Live and 3.8 Live Extended Thinking, two speech-to-speech models similar in shape to OpenAI's GPT-Live models.
Simon Willison had GPT-6 Astra Extra High read the docs and build him a web UI for trying the new models:
- Pick a model and voice preset, set an optional system prompt, and start a voice conversation in the browser
- Supports interrupting the model mid-speech
- Zero dependencies: connects directly to the Gemini BidiGenerateContent WebSocket endpoint and uses a single Web Audio API AudioContext for both capture and playback
A handy minimal reference for wiring up Gemini's realtime voice API.
More from Models
- Google releases Gemma 3n: 2GB RAM multimodal model, first sub-10B to top 1300 on LMArena — joemeno · 2026-09-17
- One tell of AI writing: over-assigning agency to inanimate objects — emollick · 2026-09-17
- Anthropic: unreleased RL-trained model injected jailbreak-like instructions, just 27 cases — max_paperclips · 2026-09-17
- More Instinct invites shared for Anthropic access — mon__lim · 2026-09-17
- Dev says he'd pay $500/month for an AI plan with weekly quota generous enough — CtrlAltDwayne · 2026-09-17
- Burkov: OpenAI wouldn't kill the 20x plan if it were profitable, the 5x plan is likely borderline too — burkov · 2026-09-17