Running Nemotron Omni Natively in Pure MLX on Mac

divinetribe1 · reddit · 2026-08-07

Nvidia's open-weights Nemotron Omni model supports vision and audio, but the existing 4-bit MLX quantization only loaded the text backbone on Mac. To fix this, the author rewrote the vision and audio towers, processor, and token splicing in pure MLX.

Original post →

More from Infra

Infra channel →