Samsung's LittleBit squeezes Llama2-13B to 0.9GB with 0.1-bit quantization
jon_durbin · x · 2026-10-08
Samsung Research introduces LittleBit, an extreme sub-1-bit LLM quantization method:
- Approach: represents weights via low-rank latent matrix factorization, then binarizes the factors, targeting 0.1 bits per weight for 31× memory reduction (Llama2-13B under 0.9GB).
- Techniques: multi-scale compensation (row, column, plus a latent dimension learning per-rank importance), Dual Sign-Value-Independent Decomposition (Dual-SVID) for stable QAT initialization, and residual compensation.
- Results: 0.1 BPW on Llama2-7B beats the prior leading method at 0.7 BPW; kernel-level benchmarks suggest up to 5× speedup vs FP16.
Related event: Samsung Open-Sources LittleBit: 13B LLM Compressed Under 1GB(4 posts)→
More from Research
- EA-VAE paper in IEEE T-PAMI fixes systematic uncertainty failures in VAEs — enzoferrante · 2026-10-10
- OpenAI reportedly solved 92 of the 500 most important open math problems in one GitHub push — altryne · 2026-10-10
- AI math is teleportation to a foggy summit: Cepelewicz's striking mountain metaphor — burny_tech · 2026-10-10
- Isola lays out the three main pushbacks to the PRH narrative his new paper answers — phillip_isola · 2026-10-10
- Three findings from Isola's new paper: global map, text-to-image translation, orthogonality — phillip_isola · 2026-10-10
- Isola's team shows a global orthogonal map aligns image and text embeddings without paired data — phillip_isola · 2026-10-10