Nemotron-3.5-Lightning runs on 16GB VRAM via ShimQuant patch
Daxfortuna · reddit · 2026-08-30
The author developed ShimQuant to optimize Nemotron-3.5-Lightning quantization. Due to tensor width misalignment with 256, standard tools failed to compress it effectively. By shimming rows during quantization and slicing at inference, the author achieved 3.07 bpw (11.77 GiB), enabling it to run on 16GB cards with a HumanEval score of 91.5%, matching larger builds. This requires a patched llama.cpp.
More from Infra
- Qwen3.8 Flash hits 415 tok/s on dual DGX Sparks — NVIDIAAI · 2026-09-01
- OpenAI's 'Jalapeno' Chip Revealed: 1500 Tokens/s Throughput — firstadopter · 2026-09-01
- TensorSharp vs llama.cpp: Qwen 3.8 Flash Next Benchmarks — fuzhongkai · 2026-09-01
- Tencent Hunyuan AngelSlim: Compressing Hy4 Model to 214GB with Heterogeneous Inference — 腾讯混元 · 2026-09-01
- Samsung shifts to 8-layer HBM4E for Nvidia with ~20% higher speed spec — 创业邦 · 2026-09-01
- Why did increasing context size increase speed in Llama.cpp? — satnl · 2026-09-01