Run a 37GB Qwen MoE in a browser tab: LocalMind streams expert weights from disk, matches llama.cpp output

naklitechie · reddit · 2026-10-06

LocalMind is a static web page (no server, no install) that runs local models via WebGPU in a browser tab. The new release streams MoE expert weights from disk during generation, letting a tab run models larger than system RAM with output that matches llama.cpp token for token.

How it works: the GGUF is copied into OPFS (the browser's private file system) on first load; dense weights, routers and KV cache live on the GPU, while routed experts stay on disk, read on demand by a pool of workers with sync access handles into a GPU slot cache (LRU, two-layer prefetch). Trunk kernels are hand-written WGSL following llama.cpp's graphs for exact parity testing.

Measured (MacBook M4 Pro 24GB, Chrome):

The whole app is still one index.html (854KB brotli), with the disk tier extracted as a standalone library (diskformer.js). The author believes this is the first browser engine reading weights from disk during generation; limits: Chrome/Edge + WebGPU only, tested on one machine, and still 3x slower than native.

Original post →

More from coding & agent

coding & agent channel →