27B Model Compressed to 3.8GB

victormustar · x · 2026-07-15

It was shared that Bonsai 27B utilizes 1-bit quantization to shrink its size from 54GB to 3.8GB, reportedly retaining 90% of its intelligence.

The post also notes that, powered by custom WebGPU kernels written using Fable 5 and GPT 5.6 Sol, the model can now run locally directly in the browser.

Related event: Bonsai 27B: The first 27B model that runs on phones(15 posts)→

Original post →

More from Infra

Infra channel →