Retraining fp16 scales only: patching a 3-bit Qwen3.8-27B GGUF closer to its BF16 parent
ZenZombie117 · reddit · 2026-09-17
A Reddit user shows how to retrain only the fp16 scale fields inside a GGUF—same bytes, same loader—to nudge ISTA-DASLab's 3-bit Qwen3.8-27B IQ3S closer to its BF16 parent. Training 103.5M blocks on 0.6M tokens with top-64 forward KL moved mean KLD from 0.0480 to 0.0450 and top-1 from 86.60% to 87.20%, though the paired 95% interval [-0.0098, +0.0038] means the honest claim is 'consistent direction, unresolved magnitude.' HellaSwag accuracy was unchanged within error, and the author stresses fidelity and benchmark equivalence are different quantities. The method is effectively EfficientQAT's second phase applied to an existing file, requires no requantisation, and the patched file plus trainer (Apache 2.0) are on Hugging Face.
More from Infra
- Qwen3.8-Flash-Next on SGLang NVFP4: 254K-Context TTFT Drops 35s→22s, Full Gauntlet Tested — FantasticNature7590 · 2026-09-17
- Privacy-focused LLM service Venice hits 250B daily tokens, up 2.5x in months — 0xAllen_ · 2026-09-17
- Baseten launches Hosted Tools, bringing server-side web search to open models — baseten · 2026-09-17
- Dev proposes predictive dynamic context caching for Claude Code — Sauers_ · 2026-09-17
- Starlink expands across Latin America: 10,000 antennas to connect 8,000 schools in Honduras alone — NicoVerderosa · 2026-09-17
- Breaking the 1.58-bit Barrier: New Paper Pushes Ternary LLMs Beyond BitNet's Limits — matt_d · 2026-09-17