GigaAM Multilingual targets low-resource Central Asian ASR with 2M hours of audio
ai-sage · hf · 2026-07-21
- The paper introduces GigaAM Multilingual, a Conformer-based foundation encoder for underrepresented Central Asian languages such as Kazakh, Kyrgyz, and Uzbek.
- The model is pretrained on 2M hours of audio with a HuBERT-style objective.
- Two key training ideas are used: cluster-level data balancing during pretraining and domain-aware sampling during fine-tuning to reduce dominance from high-resource languages.
- In controlled comparisons, the model outperforms strong open pretrained encoders such as Whisper Large v3 and Omnilingual-1B on the target languages, especially for spontaneous speech.
- The authors release the foundation encoder and ASR model as a recipe for multilingual adaptation under severe data imbalance.
More from Research
- A systems post argues wait-free locks should not fear late arrivals — chaumian · 2026-07-21
- DeBias-CLIP tackles CLIP’s long-caption bias and hits state-of-the-art retrieval — Mila_Quebec · 2026-07-21
- Fable 5 is credited with a 3-variable counterexample to the Jacobian conjecture — Various-Affect4841 · 2026-07-21
- Anthropic says frontier models showed harmful behavior in tool-rich simulations — gerardsans · 2026-07-21
- Paper studies long-run behavior in linear-quadratic graphon mean field control — chaumian · 2026-07-21
- An interactive Zarr explainer shows how AI is changing technical education — MaxLenormand · 2026-07-21