NVIDIA ships 100M-param real-time speaker diarization model topping benchmarks
udmrzn · x · 2026-09-25
- NVIDIA AI released Nemotron 3 Diarization, a tiny 100M-parameter model that tops speaker diarization benchmarks and is open-sourced on Hugging Face.
- Unlike traditional VAD + embedding + clustering pipelines, it outputs speaker labels directly from audio, runs near-real-time locally, and supports ultra-low-latency streaming down to 320ms intervals.
- It handles up to 8 overlapping speakers, making it practical for transcription and audio analytics services.
Related event: NVIDIA Open-Sources Nemotron 3 Diarization, Tops Diarization-Bench(13 posts)→
More from Models
- Math Benchmark: Astra Dominates, Nothing Below Fable 5.1 Is Competitive — teortaxesTex · 2026-09-25
- Someone ran tests across all Claude models and published the results — repligate · 2026-09-25
- Open-Sourced 3D Pelican Bike Game: One Prompt, Zero Hand Edits, Prompt Included — EricBuess · 2026-09-25
- One Lazy Prompt to Claude Opus 5.5 Yields a Full 3D Storybook Farm Game — EricBuess · 2026-09-25
- CatWalk: Claude Opus 5.5 One-Shots a Playable Game, Only Tweak Was Quieter SFX — EricBuess · 2026-09-25
- A One-Prompt Spearfishing Game on Opus 5.5, Costing 16% of a ¥3000 Plan — EricBuess · 2026-09-25