Intel's BITCOS compresses ternary LLMs to 1.485 bits per weight, boosting decode up to 27%

burny_tech · x · 2026-09-21

Intel researchers unveiled BITCOS, a compression method that pushes ternary LLMs below the conventional 1.58-bit barrier without changing a single weight:

The technique could matter increasingly for running powerful AI on PCs, smartphones, robots and edge devices.

Original post →

More from Infra

Infra channel →