DIY RCO-lite dynamic quant halves KLD vs IQ3_M with only 89MB extra on Qwen 3.5 2B

jjusko20 · reddit · 2026-10-08

A Redditor shares v1 of a "poor man's RCO" dynamic quantization method, approximating IST Austria's RCO concept—choosing precision per tensor under a byte budget by optimizing task KL on the whole model.

Pipeline:

Early results on Qwen 3.5 2B: IQ3M + imatrix is 999MB with mean KLD 0.0984; RCO-lite at 1088MB (+89MB) hits 0.0452—almost half the error—and even at 947MB (52MB smaller than IQ3M) it scores 0.0906. Author admits it likely won't beat GSQ-RCO or Unsloth Dynamic V3, but it's cheap, VRAM-frugal, and documented as a reproducible pseudo-algorithm.

Original post →

More from Infra

Infra channel →