New RCOL dynamic low-bit quantization slashes KLD error at 1-bit vs llama.cpp imatrix
jjusko20 · reddit · 2026-10-09
A Redditor released RCOL Dynamic Low-Bit Quantization (IQ1M/IQ2M/IQ3M), a modification of ISTALab's RCO algorithm that quantizes models under strict VRAM budgets.
Key points:
- An approximation algorithm using several strategies; significant KLD improvements in the 1-bit and 2-bit ranges
- Model-agnostic and doesn't require loading the FP16 teacher into VRAM
- Author's own KLD comparison table (both quants calibrated on the same wikitext imatrix) shows IQ1M divergence drastically cut at nearly the same size budget vs standard llama.cpp imatrix quants
- On Qwen 3.5 2B, the RCOL IQ1M still gets 4/5 on a test while the standard imatrix IQ1M collapses
- GGUFs published on Hugging Face (trubisky/Qwen3.5-2B-RCOL); full recipes to follow
Still proof-of-concept on a small 2B model, with plans to apply to larger models.
Related event: Community-built RCOL Dynamic Low-Bit Quantization Cuts Error in Half(2 posts)→
More from Models
- Open-Source SOTA Robotics Model Goes Head-to-Head with Closed-Source General Model — eigenron · 2026-10-09
- Latest GPT is the reverse of "biased for action", users complain — altryne · 2026-10-09
- Codex throttled to 5 tok/s as dev argues local model deployment is the only fix — lxfater · 2026-10-09
- OpenAI's math claims criticized for skipping peer review, inverting the proper process — JFPuget · 2026-10-09
- Users poke fun as Anthropic tightens guardrails on swearing at its AI — CtrlAltDwayne · 2026-10-09
- Per-token cost of Claude Pro's Opus subscription works out the same as DeepSeek Flash — teortaxesTex · 2026-10-09