NVIDIA releases 99.2M-param streaming speaker diarization model that plugs into any ASR

alexcovo_eth · x · 2026-09-24

NVIDIA's Nemotron 3 Diarization is now on ModelScope, adding live speaker attribution to existing ASR workflows without replacing the transcription model. The 99.2M-parameter, 31-layer Transformer (RoPE, Streaming Sortformer architecture) streams audio and returns speaker labels and timestamps for up to 8 speakers end-to-end, skipping separate VAD/embedding/clustering/post-processing. It pairs with Nemotron ASR, Parakeet, Canary, Whisper, etc., targets meetings, contact centers, captioning and voice agents, and runs on Ampere/Hopper/Blackwell GPUs.

Related event: NVIDIA Open-Sources Nemotron 3 Diarization, Tops Diarization-Bench(10 posts)→

Original post →

More from Models

Models channel →