NVIDIA Open-Sources VoiceChat, a Full-Duplex Speech-to-Speech Model with Tool Calling
chaumian · x · 2026-09-22
NVIDIA's NemotronLabs team released a paper on VoiceChat, an open full-duplex speech-to-speech model that stands out for supporting tool calling during real-time voice conversations.
Unlike turn-based voice assistants, the full-duplex design lets the model listen and speak simultaneously, closer to natural human conversation. Tool calling enables the model to trigger external actions within a purely voice-native interaction—a key step toward practical voice agents. The paper lists a large author team of NVIDIA speech researchers.
Related event: NVIDIA Open-Sources Full-Duplex VoiceChat with Tool Calling(2 posts)→
More from Multimodal
- SoundBoost Launches AI Music Video Generator With Full Creator Pack Exports — TheChuckTone · 2026-09-22
- eidoverse-worlds ported into Unreal: chest-lamp shadows show off engine lighting — repligate · 2026-09-22
- TTS speed vs accuracy: SpaceXAI hits 87.6% at 106 chars/sec, Kokoro fastest but weakest — ArtificialAnlys · 2026-09-22
- TTS pronunciation benchmark pricing: Gemini leads at $18.31/1M chars, Kokoro cheapest — ArtificialAnlys · 2026-09-22
- New TTS pronunciation benchmark: Gemini 3.1 Flash TTS leads at 88.1% — ArtificialAnlys · 2026-09-22
- qwen-image-2.1 photorealistic generation demo shown off — cocktailpeanut · 2026-09-22