A More Flexible Approach to Ternary Quantization

LMTLS5 · reddit · 2026-07-16

This post introduces the paper **ExTernD: Expanded-Rank Ternary Decomposition**, focusing on ternary quantization (ternary PTQ). The author's core argument is that fixing the matrix size for ternary post-training quantization is a dead end. Instead, they decompose the matrix into **two ternary matrices + an internal diagonal scaling matrix**, allowing the internal rank to be arbitrarily increased. This way, quantization accuracy can theoretically continuously approach any target, with only a slightly higher memory overhead than existing methods. The post emphasizes that this trade-off—'a bit more memory for better ternary mathematical properties'—is likely well worth it.

Related event: Ternary Decomposition as an Alternative to Quantization(2 posts)→

Original post →

More from Research

Research channel →