Running Nemotron 75B on M2 Max
These_Meaning_3883 · reddit · 2026-07-12
The author successfully ran Nemotron Puzzle 75B on a 64GB M2 Max and added native `nemotron_h_puzzle` support to `mlx-lm`. ## Benchmark Results Comparing 4-bit and 5-bit expert quantization under identical prompts, 5 random seeds, and `temp 1.0 / top_p 0.95`: - 4-bit experts: checkpoint 42.03 GiB, peak memory 49.68 GB, average 14.27 tok/s, passed 24/30 tasks - 5-bit experts: checkpoint 49.88 GiB, peak memory 58.12 GB, average 10.53 tok/s, passed 21/30 tasks - Long-context retrieval: 4-bit scored 4/5, while 5-bit scored 0/5 ## Key Takeaways - The author suspects that on a 64GB machine, 5-bit quantization operates too close to the memory limit. The growth of the KV cache and working set degrades both performance and stability. - Results for 5-bit might differ on 96GB or 128GB machines. ## Technical Details of the Adaptation - Added blockwise config and tensor remapping for Nemotron-H, matching NVIDIA's FP32 norm/router behavior. - Encountered an issue where the first layer SSM output had a cosine similarity of only 0.8832 with the reference implementation. - Root cause: NVIDIA computes `softplus(dt + dt_bias)` in BF16 before casting to FP32, whereas `mlx-lm`'s shared path casts to FP32 first. Fixing this boosted similarity to 0.999998. - BF16 output heads are mandatory: quantizing the lm_head for a 131k vocabulary to 4-bit produces repetitive garbage outputs.
More from Infra
- Local AI may pay back in 6–7 years and cut long-term costs by 30–40% — DavidLinthicum · 2026-07-21
- TSMC reportedly plans up to 10% chipmaking price hikes in 2027 — kimmonismus · 2026-07-21
- More open models and llama.cpp updates are coming, says Merve Noyan — mervenoyann · 2026-07-21
- Why adding a second LLM provider breaks more than the API surface — Ok_Extension6373 · 2026-07-21
- UK AI datacentres face backlash over heat, noise and land use — nordicinst · 2026-07-21
- Fluidstack raises $830M at $7.5B valuation as Anthropic backs a $50B compute buildout — rohanpaul_ai · 2026-07-21