Nemotron-3.5-Lightning runs on 16GB VRAM via ShimQuant patch

Daxfortuna · reddit · 2026-08-30

The author developed ShimQuant to optimize Nemotron-3.5-Lightning quantization. Due to tensor width misalignment with 256, standard tools failed to compress it effectively. By shimming rows during quantization and slicing at inference, the author achieved 3.07 bpw (11.77 GiB), enabling it to run on 16GB cards with a HumanEval score of 91.5%, matching larger builds. This requires a patched llama.cpp.

Original post →

More from Infra

Infra channel →