Nvidia Launches Nemotron-Audex-30B: Unified Audio-Text LLM

pmttyji · reddit · 2026-07-07

Nvidia has launched Nemotron-Labs-Audex-30B-A3B, expanding a 30B MoE text base (activating only 3B parameters) with audio capabilities. It integrates audio understanding, speech recognition and translation, text-to-speech, audio generation, and speech-to-speech generation while retaining its original reasoning, knowledge, and agent capabilities, supporting a context length of up to 1 million tokens. The model follows the ChatML template, supports both thinking and instruct modes, and is now openly accessible on Hugging Face.

Related event: NVIDIA Releases Open-Source Audio-Text LLM Audex-30B(9 posts)→

Original post →

More from Multimodal

Multimodal channel →