Microsoft's open-source BitNet runs a 100B LLM on a single CPU with 1.58-bit weights
JafarNajafov · x · 2026-10-10
Microsoft Research open-sourced BitNet, an inference framework that replaces 16/32-bit floats with 1.58-bit ternary weights (-1/0/+1), enabling pure integer LLM inference on CPUs.
- A 100B-parameter model runs at 5-7 tokens/sec on a single CPU, no GPU or cloud needed
- 2.37-6.17x faster than llama.cpp on x86 with 82% lower energy use; 1.37-5.07x speedup on ARM
- 16-32x memory reduction vs full-precision models
The flagship BitNet b1.58 2B4T was trained on 4 trillion tokens and benchmarks competitively with same-size full-precision models. MIT-licensed, works on macOS/Linux/Windows, enabling fully offline inference and edge deployment.
More from Infra
- $5B AI Data Center Company Pulls Listing After No Buyer Would Pay the Price — YvesMulkers · 2026-10-10
- Seth Lloyd's classic 1999 paper quantifies the physical limits of an 'ultimate laptop' — burny_tech · 2026-10-10
- vLLM Semantic Router: small-first routing cuts cost 54% while adding 16.45 accuracy points — vllm_project · 2026-10-10
- Why a Laptop Beats a Dedicated Server for Streaming Sensor Data — IgorBrigadir · 2026-10-10
- UK must not be beholden to foreign AI, says Alan Turing Institute head — nordicinst · 2026-10-10
- Neural network co-processor cards from the 1990s, brochures resurface — jensgk · 2026-10-10