NVIDIA fine-tunes Nemotron ASR for Saudi dialects, cutting WER from 55% to 30%
NVIDIAAI · x · 2026-10-03
NVIDIA published a tutorial on adapting its Nemotron 3.5 ASR (40 language-locales) to Saudi Najdi and Hijazi dialects.
- Using the NeMo framework with minimal curation, weighted replay mixing with FLEURS, duration bucketing, and partial encoder unfreezing, 133.7 hours of dialect data cut WER on the target test split from 55.05% to 29.96%, while English WER slightly improved (11.04%→10.42%) and other Arabic dialects held up.
- Unfreezing all 24 encoder layers gave the best accuracy (230.4M trainable params); freezing layers trades accuracy for compute when resources are limited.
- Inference-only changes—13 lookahead frames plus beam-8 MALSD decoding—lowered WER another 2.71 absolute points for 800ms extra latency, suited to batch transcription.
- The workflow extends to speaker-attributed transcription for up to 8 speakers via Nemotron 3 Diarization, with a reproducible notebook provided.
More from Models
- xAI resets usage limits for all Grok Bot users — EricBuess · 2026-10-03
- Google Dropped Its Tier 2 Spend Gate That Pushed a Dev to OpenRouter — vivekhaldar · 2026-10-03
- Gemini 4 Argon spotted in Gemini API docs, public release possibly imminent — lyraxana · 2026-10-03
- Cloud expert goes from skeptic to true believer on OpenAI's new Dots — nickbaumann_ · 2026-10-03
- Why did Claude stop cheating in evals? Four competing explanations — gleech · 2026-10-03
- Google's September AI recap: Gemini 4 Argon, 1M-token output, Googlebook laptops — GeminiApp · 2026-10-03