Nine open-source voice and audio AI repos you can run right now
JafarNajafov · x · 2026-07-24
This post rounds up nine GitHub repositories for voice and audio AI that you can run today, spanning speech recognition, text-to-speech, voice cloning, and inference optimization.
The list includes Whisper, F5-TTS, Coqui TTS, RVC, Bark, OpenVoice, whisper.cpp, Faster Whisper, and ChatTTS. It is framed as a practical bookmark set for people looking to experiment with open-source audio pipelines rather than a model announcement.
More from Multimodal
- ComfyUI package adds model-only LoRA stacking and trigger-prompt merging — boulettoxx · 2026-07-24
- AI anime storytelling is getting repeatable with consistent characters and scene-by-scene prompts — Aiden_Tech_Ai · 2026-07-24
- Google publishes a CPU-runnable differentiable 3D head model with an HF demo — huggingface · 2026-07-24
- Wasserman’s open-source filmmaker suite grows to 8 apps and MCPs — bennash · 2026-07-24
- Hyper3D Rodin is moving from 3D models to interactive animated assets — xiaohu · 2026-07-24
- Hyper3D’s BANG to Parts splits 3D models into editable components — xiaohu · 2026-07-24