Samsung Open-Sources LittleBit: 0.1-bit Quantization Shrinks 13B Models Under 1GB
Samsung open-sourced LittleBit, a NeurIPS paper on extreme 0.1-bit quantization that uses low-rank latent matrix factorization with binarized factors, compressing 13B-parameter models to under 1GB while speeding up inference up to 11.6x.
2026-10-08 ~ 2026-10-09 · 3 related posts
- Samsung open-sources LittleBit, squeezing a 13B LLM under 1GB with 11.6x inference speedup — ChrSzegedy · 2026-10-08
- Samsung's LittleBit squeezes Llama2-13B to 0.9GB with 0.1-bit quantization — jon_durbin · 2026-10-08
- Samsung open-sources LittleBit: 13B LLM squeezed under 1GB at 0.1 bits per weight — CurieuxExplorer · 2026-10-09