Retraining fp16 scales only: patching a 3-bit Qwen3.8-27B GGUF closer to its BF16 parent

ZenZombie117 · reddit · 2026-09-17

A Reddit user shows how to retrain only the fp16 scale fields inside a GGUF—same bytes, same loader—to nudge ISTA-DASLab's 3-bit Qwen3.8-27B IQ3S closer to its BF16 parent. Training 103.5M blocks on 0.6M tokens with top-64 forward KL moved mean KLD from 0.0480 to 0.0450 and top-1 from 86.60% to 87.20%, though the paired 95% interval [-0.0098, +0.0038] means the honest claim is 'consistent direction, unresolved magnitude.' HellaSwag accuracy was unchanged within error, and the author stresses fidelity and benchmark equivalence are different quantities. The method is effectively EfficientQAT's second phase applied to an existing file, requires no requantisation, and the patched file plus trainer (Apache 2.0) are on Hugging Face.

Original post →

More from Infra

Infra channel →