Fine-tuning Qwen3-TTS With Emotion Tags — and Emotion Vectors That Transfer Across Speakers
ProfessionalHorse707 · reddit · 2026-09-10
A developer fine-tuned Qwen3-TTS with inline transcript control tags (teacher-student distillation over 74k clips, LoRA) and open-sourced the model. Key lessons: mixed codec language prefixes fix buzzing artifacts, and high vLLM concurrency corrupts prosody. Surprisingly, emotion behaves roughly affinely in speaker-embedding space — per-emotion task vectors computed from centroids can transfer emotion control to arbitrary cloned voices without per-speaker training.
More from Multimodal
- fal keeps shipping new models and has fixed its sketchy billing, user notes — nijfranck · 2026-09-10
- The Prompt Behind That 75M-View Viral Video Is Finally Out — techhalla · 2026-09-10
- Stanford releases RenderFormer-V2: transformer neural rendering with heterogeneous scene support — Stanford · 2026-09-10
- Nitx Studio launches 'GPT Image 2.5' integration, unverified by OpenAI — aziz4ai · 2026-09-10
- SOL Attention + SageAttention Speeds Up H3 Video Gen Up to 2.5x on Apple Silicon — TgoAI · 2026-09-10
- Runway Launches MCP to Bring Video Generation into Claude, ChatGPT and Cursor — tlakomy · 2026-09-10