27B 1-bit model runs in the browser at 25-30 tok/s on a 6GB RTX 3060 laptop (WebGPU)

mentria-ai · reddit · 2026-09-09

A solo-built WebGPU/WGSL browser inference engine (mentria.ai) runs Prism ML's natively 1-bit Bonsai-27B at up to 30 tok/s decode in Chrome on a 6GB RTX 3060 laptop—no install, fully local. Key engineering: 1.14 bits/param packs 27B into 3.8GB; a 16-entry lookup table for 1-bit matvec kernels; fixing a 32-bank scratch conflict on Ampere (38% bandwidth → resolved with row padding) took decode from 15 to 32 raw tok/s; retiled prompt processing cut a 1,489-token prompt from 29.6s to 25.3s. Exact KV cache gives 3,072 context on 6GB; smaller Qwen tiers, MB-scale LoRA hot-swap and vision input are also in the engine.

Original post →

More from Infra

Infra channel →