Run a 37GB Qwen MoE in a browser tab: LocalMind streams expert weights from disk, matches llama.cpp output
naklitechie · reddit · 2026-10-06
LocalMind is a static web page (no server, no install) that runs local models via WebGPU in a browser tab. The new release streams MoE expert weights from disk during generation, letting a tab run models larger than system RAM with output that matches llama.cpp token for token.
How it works: the GGUF is copied into OPFS (the browser's private file system) on first load; dense weights, routers and KV cache live on the GPU, while routed experts stay on disk, read on demand by a pool of workers with sync access handles into a GPU slot cache (LRU, two-layer prefetch). Trunk kernels are hand-written WGSL following llama.cpp's graphs for exact parity testing.
Measured (MacBook M4 Pro 24GB, Chrome):
- Gemma 4 26B-A4B (QAT Q40, 14.4GB): character-identical to llama.cpp Metal on 9/9 conversations, 15/16 fresh prompts token-identical; 23.6 tok/s decode (3x slower than llama.cpp's 70.6); GPU process only 6.9GB, 8.6GB of experts on disk.
- Qwen3.6 35B-A3B (Q80, 36.9GB, bigger than RAM): GPU process 7.3GB, 9.9 tok/s decode; per-token cost 23ms GPU, 39ms routing round trips, 35ms SSD expert reads; moving routing to the GPU gave no gain.
- Gemma 4 E2B's 1.2GB per-layer embedding table can also stay on disk: 4.27→2.07GB GPU memory, identical output, 3-8% slower.
The whole app is still one index.html (854KB brotli), with the disk tier extracted as a standalone library (diskformer.js). The author believes this is the first browser engine reading weights from disk during generation; limits: Chrome/Edge + WebGPU only, tested on one machine, and still 3x slower than native.
More from coding & agent
- OpenAI streamlines ChatGPT plugin submissions: upload zip, fix validation, publish — Dimillian · 2026-10-06
- Long-running benchmarks find Strata inference server failing full-build scenarios — julianharris · 2026-10-06
- OpenAI makes Codex Auto-review free for all: a second agent vets high-risk actions — kimmonismus · 2026-10-06
- SourceLearn Builds Source-Specific Agent Competence, Wins 13 of 15 Benchmarks — GeorgiaTech · 2026-10-06
- Programmatic Search Agents boost task success by up to 7.56 points over query-based agents — _reachsumit · 2026-10-06
- SWE-Race: 188 real concurrency bugs benchmark where GPT-5.6 Luna scores 81% — heyitsdannyle · 2026-10-06