Best Quantization for Sub-2-bit? Beyond QTIP, What Papers to Read?

Aggravating-Push-207 · reddit · 2026-08-12

A user planning to implement their own inference engine for very large LLMs (e.g., Qwen 3.6/8) with only 8GB VRAM is researching sub-2-bit quantization. They found QTIP as the best so far but seek the latest research and paper recommendations.

Original post →

More from Infra

Infra channel →