Intel's BITCOS compresses ternary LLMs to 1.485 bits per weight, boosting decode up to 27%
burny_tech · x · 2026-09-21
Intel researchers unveiled BITCOS, a compression method that pushes ternary LLMs below the conventional 1.58-bit barrier without changing a single weight:
- It exploits the unusually high fraction of zero weights in ternary models — across 29 ternary LLM checkpoints, zeros reached up to 51.48%
- This allows BITCOS to reach just 1.485 bits per weight while preserving exact ternary weights
- In testing it delivered up to 18% higher CPU decode throughput and up to 27% higher GPU decode throughput: smaller models, less memory movement, faster inference.
The technique could matter increasingly for running powerful AI on PCs, smartphones, robots and edge devices.
More from Infra
- Laya, an open-source local AI, is being built into Omarchy M to run on Apple GPUs and ANE — Scobleizer · 2026-09-21
- Facebook and Instagram down for thousands of users in ongoing outage — Polymarket · 2026-09-21
- Agent builders say provider KV caching black boxes block swarm and long-run agents — Small_Luck8177 · 2026-09-21
- Laya ported to MLX: M3 Max runs local agent 60 decisions/sec, 50x faster — sven_ai · 2026-09-21
- US Data Center and Info-Processing Hardware Spending Now Exceeds Housing Investment — rohanpaul_ai · 2026-09-21
- iPhone 18 Pro runs 27B models 2x faster; tease of 100B+ local LLM on iPhone — MannyKayy · 2026-09-21