Llama.wasm: llama.cpp Compiled to WASM with WebGPU Runs LLMs Fully in the Browser
Numerous-Fan8138 · reddit · 2026-10-06
The author open-sourced Llama.wasm, a technical proof of concept showing an LLM running entirely client-side in a browser using llama.cpp compiled to WASM with WebGPU.
Key points:
- llama.cpp compiled to WebAssembly, runnable directly in any modern browser
- WebGPU provides on-device GPU-accelerated inference
- A new path for local deployment: no install, no backend, user data never leaves the device
This is the repo behind the demo video posted separately by the same author.
Related event: llama.cpp runs LLMs fully in-browser via WASM and WebGPU(3 posts)→
More from Infra
- Learning electronics with Opus: two weeks of experiments distilled into interactive ET-SoC-1 diagrams — yaroslavvb · 2026-10-06
- Claude Code arrives in AWS GovCloud, bringing AI coding to ITAR-regulated workloads — AWS ML Blog · 2026-10-06
- The AI Stack Now Extends to the Power Plant as Google, Amazon, Meta Chase Nuclear — ingliguori · 2026-10-06
- LithosAI Launches LithosBox Millisecond Agent Sandboxes; Hits 727 TPS on GLM 5.3 Flash — JiaZhihao · 2026-10-06
- LithosAI Claims Third #1 Speed Spot: Fastest Inference for GLM 5.3 Flash on Artificial Analysis — JiaZhihao · 2026-10-06
- AWS ships aws-ai-ml skill that turns coding agents into SageMaker inference optimization experts — AWS ML Blog · 2026-10-06