Fully Open-Source Real-Time Voice Digital Human Pipeline

victormustar · x · 2026-07-03

Developer victormustar shared a fully open-source real-time voice digital human pipeline: Silero VAD for voice activity detection, Parakeet for speech recognition, Gemma 4 31B (running on Cerebras for fast responses) for dialogue, and Qwen3-TTS for speech synthesis, transmitting PCM audio via native WebSocket. The digital human's avatar and lip-sync use the TalkingHead and HeadAudio projects. The entire solution is built exclusively on open models.

Original post →

More from Infra

Infra channel →