Gemini Audio Launch Party: Devs Build Steerable TTS Apps, Sushi Translation Demo Steals Show
jocarrasqueira · x · 2026-09-25
At the Gemini audio launch party, founders and devs showed off what they're building with Gemini TTS: natural, expressive voices steerable via simple prompts controlling tone, pace, emotion, and even multiple speakers. The favorite demo was a sushi station that translates your order into any language—a use case the author sees fitting any client-facing business.
Related event: Gemini Audio Event Wows with Real-Time English-Japanese Sushi Ordering(3 posts)→
More from Multimodal
- FineVision, the 17M-image open VLM dataset from 200+ sources, accepted to NeurIPS — andimarafioti · 2026-09-25
- MiniMax H3 seems overtrained on smiles: 'bored caterpillar' video prompt keeps breaking immersion — episodex86 · 2026-09-25
- Pose Blueprint: A Browser-Based 3D Pose Editor for ComfyUI and ControlNet — OkConfusion6667 · 2026-09-25
- Reddit User Explores AI Art With Only Steps, CFG and Denoise Tweaks, No LoRAs — Extreme_Nice · 2026-09-25
- A sub-$20 LoRA makes Qwen-Image 2.1 rotate transparent objects with a prompt — ben_burtenshaw · 2026-09-25
- Lingbot World v2 runs at 60 FPS, hinting world models could reshape game dev — bingxu_ · 2026-09-25