DIY RCO-lite dynamic quant halves KLD vs IQ3_M with only 89MB extra on Qwen 3.5 2B
jjusko20 · reddit · 2026-10-08
A Redditor shares v1 of a "poor man's RCO" dynamic quantization method, approximating IST Austria's RCO concept—choosing precision per tensor under a byte budget by optimizing task KL on the whole model.
Pipeline:
- Build an imatrix "sensitivity curve" per tensor on the full-precision model to extrapolate which tensors suffer most from quantization.
- Iterative local search substitutes precisions under a size cap, discarding any change that raises KL divergence.
- The teacher model loads once to dump token logprobs (CPU-friendly); afterward you only need VRAM for the target size +20-25% buffer. Runs take minutes to hours.
Early results on Qwen 3.5 2B: IQ3M + imatrix is 999MB with mean KLD 0.0984; RCO-lite at 1088MB (+89MB) hits 0.0452—almost half the error—and even at 947MB (52MB smaller than IQ3M) it scores 0.0906. Author admits it likely won't beat GSQ-RCO or Unsloth Dynamic V3, but it's cheap, VRAM-frugal, and documented as a reproducible pseudo-algorithm.
More from Infra
- Reliquary opens its post-training network to Bittensor subnet teams, starting with batch data generation — const_reborn · 2026-10-08
- Samsung Q3: watch whether sales clear ₩200T — it would mean memory pricing runs hotter than contracts show — tengyanAI · 2026-10-08
- Google's Project Suncatcher starts testing TPUs in orbit for scalable AI compute — ZoubinGhahrama1 · 2026-10-08
- Claude Haiku 5.5 hits Databricks Day 0: ~15% better than Haiku 4.5 at a fraction of the cost — matei_zaharia · 2026-10-08
- Cloudflare ships cf sql query: SQL over all Cloudflare analytics in the terminal — irvinebroque · 2026-10-08
- Using AI-assisted math to fit an 11-Mac-mini home cluster on the smallest table — BLUECOW009 · 2026-10-08