voice-pro: A Gradio WebUI Integrating Zero-Shot Voice Cloning and Multilingual Translation
abus-aikorea · github · 2026-08-01
The open-source project voice-pro provides a one-stop Gradio WebUI audio processing solution for creators and developers.
It integrates text-to-speech (TTS) tools like Edge-TTS and kokoro, and supports zero-shot voice cloning using E2 & F5-TTS and CosyVoice. Additionally, it features Whisper speech recognition, YouTube downloading, Demucs vocal isolation, and multilingual translation.
More from Multimodal
- Training a Krea 2 LoRA for amateur smartphone-style photos with Indian vibe — desiNaughtyAI · 2026-08-01
- ByteDance uses Seedance 2.5 to push 'cinematic' lessons in study app — JayaIsNotGemez · 2026-08-01
- Midjourney Style Reference Code Powers Surreal AI Video Workflow — azed_ai · 2026-08-01
- xAI Upgrades Imagine Video 1.5: Adds Text-to-Video, Native 1080p, and Multi-References — testingcatalog · 2026-08-01
- Microsoft Open-Sources TRELLIS.2: 3D Generation with Compact Structured Latents — microsoft · 2026-08-01
- Solo Creator Produces 46-Minute AI Sci-Fi Film, Shares Workflow — ABePick · 2026-08-01