DEFINE Decouples Accent from Voice in Zero-Shot TTS, Matching Two-Model Cascades with One Model
amaai-lab · hf · 2026-10-05
amaai-lab presents DEFINE, an end-to-end zero-shot TTS framework that decouples speaker identity from target accent, each conditioned on separate audio exemplars.
Highlights
- Built on F5-TTS with LoRA adaptation; an exemplar encoder maps short accent clips into conditioning space via learned accent prototypes.
- No accent labels at inference, no waveform post-conversion; a single guidance weight controls accent strength without retraining.
- On seen accents, stronger guidance lifts accent-probe accuracy from 6.5% to 19.6%.
- One model matches a two-model TTS-voice-conversion cascade on seen and out-of-domain accents with higher speaker similarity; not on held-out accents.
More from Multimodal
- One Prompt Gets Fable to Generate a 5,000-Year History of China Video — FuSheng_0306 · 2026-10-05
- UniMate Releases 3D Rigged Skeleton Models on Hugging Face, ComfyUI Integration Proposed — RazsterOxzine · 2026-10-05
- Full Making-of Released for AI Music Video 'Close All the Windows' — All Prompts and Orchestration Logs Included — Afinetheorem · 2026-10-05
- OctLLM encodes 3D geometry as explicit octree token sequences without sacrificing language ability — _akhaliq · 2026-10-05
- Gemma 31b 'Grand Horror' + H3 generates atmospheric horror video — jrexthrilla · 2026-10-05
- EditHero: First benchmark for long-horizon part-level 3D editing compares agentic vs non-agentic approaches — _akhaliq · 2026-10-05