Zero-Dependency C99 BitNet Inference Hits 36 tok/s on Intel Xeon CPU
shifu_legend · reddit · 2026-08-09
A developer built a CPU-first inference engine from scratch in pure C99, running BitNet 1.58-bit ternary models natively without Python, CUDA, or BLAS.
Core Tech & Performance:
- Achieves 36.25 tok/s on BitNet b1.58-2B-4T using 4 threads on an Intel Xeon.
- Native ternary SIMD: Weights are packed 4 per byte; custom AVX2/AVX-512 routines with VNNI instructions accumulate directly into integer registers, avoiding float32 unpacking overhead.
- Minimal runtime: Thread pool uses C11 atomics with spin-then-yield backoff for near-zero sync overhead. Compiles to a standalone binary serving an OpenAI-compatible API.
Bottleneck & Takeaways:
- Hits the DRAM bandwidth ceiling. At batch size 1, decode speed is memory-bound. Running at 95% of theoretical memory bandwidth, further compute kernel optimizations won't improve end-to-end latency.
More from Infra
- AI Data Centers End Decades of Stagnant US Power Use, Challenging Anti-Growth Mindsets — AndyMasley · 2026-08-09
- Amazon's New Texas Data Center Power Plant Could Become a Top US Polluter — The Verge AI · 2026-08-09
- Upgrading to PyTorch 2.13 and CUDA 13 Doubles Video Resolution on Blackwell GPUs — Chemical-Bicycle3240 · 2026-08-09
- Training 10T Parameter Models Requires Over 50k GB300 GPUs — AccBalanced · 2026-08-09
- Amazon Data Centers Become the Biggest Pollution Source in the US — geox · 2026-08-09
- Matt Turck: Tech Industry Shouldn't Ignore Resistance to AI Data Centers — mattturck · 2026-08-09