Base-3 packing for ternary GGUFs: ~22% less weight VRAM, lossless
pmttyji · reddit · 2026-09-05
Developer llopresto87 built a denser GGUF format Q2B3 / B3S for ternary models (BitNet-b1.58, Ternary-Bonsai): since weights are only {-1, 0, +1} times a block scale, they pack directly in base 3 — 128 weights need just 26 bytes of trits plus one f16 scale, i.e. 1.75 bits per weight.
Weight sizes:
- 9B: 2.5 GB Q20 → 2.0 GB B3S
- 27B: 7.6 GB Q20 → 5.9 GB B3S
Key points:
- Lossless for genuinely ternary models — same states stored, just base-3 packing instead of a general 2-bit scheme; feeding it an FP16 model will destroy quality.
- Implemented as a small llama.cpp fork (commit 4e97ac86e); main path is AMD ROCm, tuned on a 7900 XTX; CPU works; CUDA/Metal compile but are unverified.
- No perceptible decoding speed loss on the author's hardware; full speed/perplexity tables pending.
- A separate repacker converts older 30-byte/dual-scale blocks to the new layout, verifying each block and aborting rather than silently producing a lossy file.
Most useful help needed: NVIDIA/Apple users verifying CUDA/Metal output vs CPU. Repos: llama-cpp-ternary-b3s and ternary-q20-repacker.
More from Infra
- Hugging Face downloads crawl at 700kb/s on 1Gbit lines, users float p2p model distribution — perelmanych · 2026-09-05
- After ChatGPT, Claude & Grok All Went Dark, One User's 3-Machine Local AI Lab Kept Running — cocktailpeanut · 2026-09-05
- DRAM density has flattened: servers stuck at 8TB for 5 years, CXL is the way out — lauriewired · 2026-09-05
- Benchmarked 21 Qwen3.8-27B quants on 16GB VRAM: bartowski IQ4_XS wins — Storterald · 2026-09-05
- Nvidia DLSS 5 frame interpolation discussed in Stable Diffusion community — KonoTheSavage1 · 2026-09-05
- Dylan Patel on Dwarkesh: How Elon Musk Played the Compute Market — Dwarkesh Patel · 2026-09-05