Samsung open-sources LittleBit: 13B LLM squeezed under 1GB at 0.1 bits per weight
CurieuxExplorer · x · 2026-10-09
Samsung open-sourced LittleBit, a NeurIPS method that compresses a 13B-parameter LLM to under 1GB.
- Instead of storing weights as numbers, it uses latent factorization to push them to extreme sub-1-bit levels — down to 0.1 bits per weight in some configs, beating 0.7-bit methods.
- Architecturally radical: at this compression level floating-point multiplication is replaced by bitwise XOR — sign flips instead of matrix math.
- Reported 11.6x inference speedup; the authors claim "the memory wall just moved."
More from Infra
- Oxide Computer raises $445M Series D as AI demand drives enterprises to rethink owning vs renting compute — Sethwinterroth · 2026-10-09
- Datology AI launches Curation Studio, pitching data quality as the ultimate compute multiplier — schwarzjn_ · 2026-10-09
- Impactful Scheduling for GPU Clusters: Inside AI2's New Scheduler — Hugging Face Blog · 2026-10-09
- The 'Montreal Premium': Same GPUs Cost 20%+ More Per Hour in Canada Than the US — RichardsonDx · 2026-10-09
- Supabase acquires Turso and rebuilds itself around coding agents at Select 2026 — glcst · 2026-10-09
- Epoch AI: AI costs falling 47%/quarter, 54x faster than electricity's decline — Jsevillamol · 2026-10-09