NVIDIA Releases Nemotron 3 Diarization, Massively Outperforming Rivals
Sam Witteveen · youtube · 2026-09-23
Sam Witteveen reviews NVIDIA's new Nemotron 3 Diarization model for speaker diarization, supporting both batch and streaming modes, which he says massively outperforms other models at establishing who said what.
The video covers the Nemotron speech family, what diarization is and how models are scored (DER), a breakdown of the model, and hands-on demos on DGX Spark with NeMo — including processing a full podcast and exporting to text or SRT. The model is open on Hugging Face.
Related event: NVIDIA Open-Sources Nemotron 3 Diarization, Tops Diarization-Bench(9 posts)→
More from Multimodal
- Viggle's distilled Qwen-Image 2.1 turbo LoRA trends on Hugging Face — Viggle · 2026-09-24
- Google launches Gemini 3.8 Flash TTS with custom voices in 100+ languages — kastnerkyle · 2026-09-24
- jxnl's robust music transcription pipeline: Astra + spectrogram error correction — jxnlco · 2026-09-24
- AI-made music video on datacenter water use: GPT 6 + Suno + Slop Cannon pipeline — zealcaiden · 2026-09-24
- AI Demo Generates Explorable Mars Life, Takes Real-Time Scene Direction — karinanguyen · 2026-09-24
- ACTx486 demo lets you talk to any video and get real-time responses — karinanguyen · 2026-09-24