27B 1-bit model runs in the browser at 25-30 tok/s on a 6GB RTX 3060 laptop (WebGPU)
mentria-ai · reddit · 2026-09-09
A solo-built WebGPU/WGSL browser inference engine (mentria.ai) runs Prism ML's natively 1-bit Bonsai-27B at up to 30 tok/s decode in Chrome on a 6GB RTX 3060 laptop—no install, fully local. Key engineering: 1.14 bits/param packs 27B into 3.8GB; a 16-entry lookup table for 1-bit matvec kernels; fixing a 32-bank scratch conflict on Ampere (38% bandwidth → resolved with row padding) took decode from 15 to 32 raw tok/s; retiled prompt processing cut a 1,489-token prompt from 29.6s to 25.3s. Exact KV cache gives 3,072 context on 6GB; smaller Qwen tiers, MB-scale LoRA hot-swap and vision input are also in the engine.
More from Infra
- vLLM x AgentX: Full-Stack Optimizations for Real-World Agentic Serving — jfiance · 2026-09-09
- Desert Ant Labs introduces on-device intelligence for every product — Arcuru · 2026-09-09
- Speculative Decoding With Qwen3-30B-A3B Yields 1.5x Local Speedup, Up to 5x — Arindam_1729 · 2026-09-09
- Explainer: Speculative Decoding Speeds Up LLM Inference by ~100% — blaizedsouza · 2026-09-09
- Cerebras CTO's chip architecture deep dives—WSE-3, Hot Chips 34, Cornell lectures—barely get any views — blaizedsouza · 2026-09-09
- Cosmos3 (64B) INT4 Quants Bring Local Image and Video Gen to Mac and CUDA — Formal-Swordfish-228 · 2026-09-09