WordVoice TTS: Per-Word Control of Duration, Loudness, Pitch, Tone in a 0.5B Model
multimodalart · x · 2026-07-24
WordVoice TTS launches with per-word control over duration, loudness, pitch, and tone. Built on CosyVoice3, it offers auto-pilot or manual control, voice cloning, and preset voices. The 0.5B parameter model performs well for its size.
More from Multimodal
- Opus 5 is being praised for 3D output, but users say it is still slow and token-hungry — OfirPress · 2026-07-24
- Sonilo says it can generate licensed sound effects from video in seconds — socialwithaayan · 2026-07-24
- Sonilo pitches video-aware sound effects with no timeline editing — socialwithaayan · 2026-07-24
- Reddit explores a Blender-first pipeline for consistent AI video keyframes — Alone-Performer5065 · 2026-07-24
- Runway users say audio prompts are optional and video timing sets the beat — iamfakhrealam · 2026-07-24
- Sonilo launches Sound Effects v1.0 to auto-place realistic audio on videos — iamfakhrealam · 2026-07-24