Base-3 packing for ternary GGUFs: ~22% less weight VRAM, lossless

pmttyji · reddit · 2026-09-05

Developer llopresto87 built a denser GGUF format Q2B3 / B3S for ternary models (BitNet-b1.58, Ternary-Bonsai): since weights are only {-1, 0, +1} times a block scale, they pack directly in base 3 — 128 weights need just 26 bytes of trits plus one f16 scale, i.e. 1.75 bits per weight.

Weight sizes:

Key points:

Most useful help needed: NVIDIA/Apple users verifying CUDA/Metal output vs CPU. Repos: llama-cpp-ternary-b3s and ternary-q20-repacker.

Original post →

More from Infra

Infra channel →