NVIDIA Releases Open-Source Audio-Text LLM Audex-30B

NVIDIA has officially released Nemotron-Labs-Audex-30B-A3B (Audex), an open-source unified audio-text Large Language Model. Designed to process both audio and text simultaneously within a single architecture, the model addresses the intelligent bottlenecks of traditional audio models in multimodal fusion, marking a significant advancement in multimodal AI.

Architecture and Core Capabilities

Audex utilizes a Mixture-of-Experts (MoE) architecture with a total of 30B parameters but only activates 3B parameters during inference. By sharing a transformer decoder, the model projects audio inputs into the text embedding space to integrate audio and text processing. This design allows it to natively support full-stack tasks including audio understanding, speech recognition and translation, text-to-speech, audio generation, and speech-to-speech generation, while supporting SFT and RL training.

Performance and Industry Reception

NVIDIA emphasizes that Audex achieves "audio intelligence without sacrificing text intelligence." It achieved state-of-the-art performance among open-source models on multiple audio intelligence benchmarks, while also outperforming text-only models in math, code reasoning, alignment, and instruction following. Developer @andrew_n_carr noted that while many audio models lack common sense, Audex demonstrates stronger comprehension, potentially making it the first native audio model that is "not so dumb."

2026-07-07 ~ 2026-07-09 · 9 related posts

2 near-duplicate retellings: _weiping · _weiping