Nvidia Launches Nemotron-Audex-30B: Unified Audio-Text LLM
pmttyji · reddit · 2026-07-07
Nvidia has launched Nemotron-Labs-Audex-30B-A3B, expanding a 30B MoE text base (activating only 3B parameters) with audio capabilities. It integrates audio understanding, speech recognition and translation, text-to-speech, audio generation, and speech-to-speech generation while retaining its original reasoning, knowledge, and agent capabilities, supporting a context length of up to 1 million tokens. The model follows the ChatML template, supports both thinking and instruct modes, and is now openly accessible on Hugging Face.
Related event: NVIDIA Releases Open-Source Audio-Text LLM Audex-30B(9 posts)→
More from Multimodal
- A fine-tuned Krea 2 raw model produced a rainy-night driving scene — darlens13 · 2026-07-27
- Users ask whether Video2X can load custom OpenModelDB models — Used-Profit2355 · 2026-07-27
- A builder wants AI to reverse-engineer viral video effects into ComfyUI workflows — stale2000 · 2026-07-27
- Midjourney V8.2 adds personalization and shows off stylized image outputs — Mr_AllenT · 2026-07-27
- Midjourney’s image variety draws a Krea 2 comparison and asks how to reproduce it — diffusion_throwaway · 2026-07-27
- AI short film sets a 1985 dystopia to music and leans into cinema — ProfessorKey98 · 2026-07-27