OmniVoice open-sources diffusion-based TTS with zero-shot voice cloning in 600+ languages
tom_doerr · x · 2026-09-04
k2-fsa has open-sourced OmniVoice, a massively multilingual zero-shot TTS model supporting over 600 languages — claimed to be the broadest coverage among zero-shot TTS models. It already has 9.7k GitHub stars.
- Architecture: Built on a novel diffusion language model-style architecture, generating high-quality speech with fast inference
- Capabilities: State-of-the-art voice cloning plus voice design via speaker attributes (gender, age, pitch, dialect/accent, whisper)
- Fine-grained control: Non-verbal tokens like [laughter] and pronunciation correction via pinyin or phonemes
- Ships with Python API, CLI tools, and training/evaluation pipelines
More from Multimodal
- Nunchux launches: one API aggregating 30+ image, video and avatar models — junyanz89 · 2026-09-04
- User combines GPT-6 Astra with a modified fal H3 Max Director in new video prototype — OdinLovis · 2026-09-04
- Ideogram's Painful bbox/ Prompting Made the Minimax-H3 Transition Easy — Nimblecloud13 · 2026-09-04
- Higgsfield demos text-to-3D pipeline: GPT-6 Astra vibes-codes an Oval Office scene in Blender — jxnlco · 2026-09-04
- First Video Tutorial: Building Character Sheets in Krea 2 for MiniMax-H3 — solomars3 · 2026-09-04
- Stolen Texture: a GPT Image 2 prompt that spreads product material to everything it touches — aziz4ai · 2026-09-04