Google Gemini 3.8 Flash TTS heard in examples: shorthand expansion and contextual pronunciation
ArtificialAnlys · x · 2026-09-24
Artificial Analysis shared example audio reportedly generated by Gemini 3.8 Flash TTS, showcasing two speech generation challenges:
- Expanding shorthand: e.g. "Median latency was 12 ms, reported as p50." (ms, p50 read in full) and "Take the elevator to the 3rd Fl and turn left." (Fl → Floor).
- Contextually appropriate disambiguation: e.g. "St. Mary's is on Church St." (St. read as Saint vs Street).
More from Multimodal
- Claude Opus 5.5 generates a music video in ~1 shot: "not good, but not without interest" — NathanpmYoung · 2026-09-24
- Gemini 3.8 Flash TTS and Flash-Lite TTS land on Merge Gateway, top Hume voice quality index — shensi · 2026-09-24
- LemonSlice Launches Character World Model-1, a Real-Time Interactive Avatar Model — mhdfaran · 2026-09-24
- Reverse workflow: unpack a reference image with Extract Prompt, then remix it — JaynitMakwana · 2026-09-24
- Unverified claim: 'GPT-6 Astra' builds full video projects via Codex + Dreamina CLI — JaynitMakwana · 2026-09-24
- MiniMax H3 roundup: video VAE 2.2x faster encoding, music model under 12GB — optimisticalish · 2026-09-24