PrismML runs a 2B vision model with 1-bit weights on 4GB smart glasses at 2x speed

lmoroney · x · 2026-09-26

Laurence Moroney breaks down several Snapdragon Summit announcements that all point the same way: tiny models quantized below 4 bits per weight are being tuned for NPUs inside glasses and phones, and the chips are being built with those models in mind.

The standout is PrismML's demo: a 2B-parameter vision-language model running fully locally on smart glasses built on Snapdragon AR1 Gen 1 — a 1.7B 1-bit language model paired with a 0.3B vision encoder at 4 bits. Qualcomm's tests on a 4GB RAM, 6 TOPS, 1024-token-context platform measured 0.43 GB of weights for the 1-bit version versus 1.66 GB for the 4-bit equivalent (3.8× smaller), at about 2× the speed of 4-bit.

The article also explains key terms (NPUs such as Qualcomm's Hexagon, quantization, prefill vs. decode), covers related work from Liquid AI and new Snapdragon chips, and lays out what builders should watch next.

Original post →

More from Embodied

Embodied channel →