NVIDIA releases 99.2M-param streaming speaker diarization model that plugs into any ASR
alexcovo_eth · x · 2026-09-24
NVIDIA's Nemotron 3 Diarization is now on ModelScope, adding live speaker attribution to existing ASR workflows without replacing the transcription model. The 99.2M-parameter, 31-layer Transformer (RoPE, Streaming Sortformer architecture) streams audio and returns speaker labels and timestamps for up to 8 speakers end-to-end, skipping separate VAD/embedding/clustering/post-processing. It pairs with Nemotron ASR, Parakeet, Canary, Whisper, etc., targets meetings, contact centers, captioning and voice agents, and runs on Ampere/Hopper/Blackwell GPUs.
Related event: NVIDIA Open-Sources Nemotron 3 Diarization, Tops Diarization-Bench(10 posts)→
More from Models
- Xiaomi ships open-weights MiMo-V2.6-Pro, tops open-source AI index at 46, rivaling Claude Opus 5 — ycombinator · 2026-09-24
- Friends in rural Sweden who only knew ChatGPT are now asking about Meta's Muse — gabriel1 · 2026-09-24
- User Teases That OpenAI's Post-Training Team Is Cooking an 'Opus 4-5' Rival — sloppenheimer · 2026-09-24
- Claude Opus 5.5 tops VoxelBench, GPT-6 Sol ranks third — legit_api · 2026-09-24
- GPT 6 computer use impresses: drives user through 15 banking sites in one go — altryne · 2026-09-24
- IKEA Assembly Benchmark: Top Model Score Jumped From 28% to 80% in 10 Months — emollick · 2026-09-24