EmoRES: training-free vector steering boosts emotion TTS controllability by up to 118.8%
meta · hf · 2026-09-30
EmoRES is a training-free vector steering method for emotional TTS. It decomposes each emotion steering vector into a shared component moving speech away from neutral and a residual component directing generation toward the requested emotion, controlling each separately without retraining the backbone. On IEMOCAP with IndexTTS-2 and CosyVoice2, rank correlation improves by 26.13 and 12.97 points (relative gains of 118.8% and 33.1%), emotion hit rate rises by up to 12.95 points, and human evaluation shows up to 35.0% relative improvement in perceived emotion and 63.8% pairwise naturalness preference.
More from Multimodal
- Four imaginary tokens for Midjourney v8.2 produce memory ghosts and bone echoes — LudovicCreator · 2026-09-30
- Hyper-personalized music is BS: music is culture and inherently social, argues developer — jordiponsdotme · 2026-09-30
- NUS Proposes StoryEngine: A State-Grounded Agentic Framework for Coherent Long-Form Video Storytelling — NationalUniversityofSingapore · 2026-09-30
- One Year of Local Image Generation: Why Civitai and ComfyUI Both Fall Short — BenDLH · 2026-09-30
- Opus made a launch video for Violetto 1B in 50 minutes amid zero media coverage — tensorqt · 2026-09-30
- LoRA adapters break on distilled video models, long post explains why — burkov · 2026-09-30