Fine-tuning Qwen3-TTS With Emotion Tags — and Emotion Vectors That Transfer Across Speakers

ProfessionalHorse707 · reddit · 2026-09-10

A developer fine-tuned Qwen3-TTS with inline transcript control tags (teacher-student distillation over 74k clips, LoRA) and open-sourced the model. Key lessons: mixed codec language prefixes fix buzzing artifacts, and high vLLM concurrency corrupts prosody. Surprisingly, emotion behaves roughly affinely in speaker-embedding space — per-emotion task vectors computed from centroids can transfer emotion control to arbitrary cloned voices without per-speaker training.

Original post →

More from Multimodal

Multimodal channel →