NVIDIA Releases Nemotron 3.5 ASR: Ultra-Low Latency but Misses Drug Names in Clinical Test

MaziyarPanahi · x · 2026-08-07

NVIDIA has released the Nemotron 3.5 ASR streaming multilingual speech recognition model (0.6B parameters) on Hugging Face.

A developer tested the model in a clinical scenario: it generated the first text chunk in just 1.1 seconds with a final lag of only 40ms, successfully capturing phrases like "five milligrams twice daily." However, it failed to recognize specific drug names such as "warfarin" and "apixaban."

This highlights a classic tradeoff in streaming speech recognition for healthcare: achieving ultra-fast response times often compromises accuracy in specialized vocabulary, suggesting a need for domain-specific fine-tuning.

Related event: NVIDIA Launches Nemotron 3.5 ASR: Ultra-Low Latency but Prone to Medical Errors(3 posts)→

Original post →

More from Models

Models channel →