EmoRES: training-free vector steering boosts emotion TTS controllability by up to 118.8%

meta · hf · 2026-09-30

EmoRES is a training-free vector steering method for emotional TTS. It decomposes each emotion steering vector into a shared component moving speech away from neutral and a residual component directing generation toward the requested emotion, controlling each separately without retraining the backbone. On IEMOCAP with IndexTTS-2 and CosyVoice2, rank correlation improves by 26.13 and 12.97 points (relative gains of 118.8% and 33.1%), emotion hit rate rises by up to 12.95 points, and human evaluation shows up to 35.0% relative improvement in perceived emotion and 63.8% pairwise naturalness preference.

Original post →

More from Multimodal

Multimodal channel →