Samsung Open-Sources LittleBit: 0.1-bit Quantization Shrinks 13B Models Under 1GB

Samsung open-sourced LittleBit, a NeurIPS paper on extreme 0.1-bit quantization that uses low-rank latent matrix factorization with binarized factors, compressing 13B-parameter models to under 1GB while speeding up inference up to 11.6x.

2026-10-08 ~ 2026-10-09 · 3 related posts