What does "expressiveness" in voice AI actually mean? A deep dive with Inworld's TTS-2
SIGKITTEN · x · 2026-09-03
"Expressiveness" is voice AI's current buzzword—this video unpacks what it actually means, using material from Inworld's new TTS-2 model. It covers the Blizzard Challenge history, naturalness vs expressiveness, human evals and speech arenas, AI-as-judge, deterministic proxies, Goodhart's law, text markers and alignment post-training for expressive models, and why expressiveness isn't always the goal.
More from Multimodal
- ComfyUI to announce MiniMax H3 sync sound challenge winners in special livestream — MiniMax_AI · 2026-09-03
- Wan 2.2 Image-to-Video Tested in ComfyUI With Two-Stage High-to-Low Noise KSampler Pass — VictorVisuals · 2026-09-03
- AI influencers reshape content creation as Seedance 2.5 targets influencer vlogs — aftahi_ai · 2026-09-03
- Seedance 2.5 tested for AI influencer vlogs: 30-second clips with consistent identity — aftahi_ai · 2026-09-03
- A crowned, sneering AI portrait in fluorescent pink and yellow — misovalko · 2026-09-03
- Oxford & NUS Propose V-RAE: Replacing Pixel Reconstruction with Pretrained Visual Representations for Video Generation — jiqizhixin · 2026-09-03