Running LLMs Fully Client-Side: llama.cpp Compiled to WASM with WebGPU Proof of Concept
Numerous-Fan8138 · reddit · 2026-10-06
The author shows a weekend experiment: a proof-of-concept running an LLM entirely client-side in the browser, by compiling llama.cpp to WASM with WebGPU. A demo video is included. All inference happens on the user's device with no server involved — a hands-on exploration of client-side/local LLM deployment.
Related event: llama.cpp runs LLMs fully in-browser via WASM and WebGPU(3 posts)→
More from Infra
- Learning electronics with Opus: two weeks of experiments distilled into interactive ET-SoC-1 diagrams — yaroslavvb · 2026-10-06
- Claude Code arrives in AWS GovCloud, bringing AI coding to ITAR-regulated workloads — AWS ML Blog · 2026-10-06
- The AI Stack Now Extends to the Power Plant as Google, Amazon, Meta Chase Nuclear — ingliguori · 2026-10-06
- LithosAI Launches LithosBox Millisecond Agent Sandboxes; Hits 727 TPS on GLM 5.3 Flash — JiaZhihao · 2026-10-06
- LithosAI Claims Third #1 Speed Spot: Fastest Inference for GLM 5.3 Flash on Artificial Analysis — JiaZhihao · 2026-10-06
- AWS ships aws-ai-ml skill that turns coding agents into SageMaker inference optimization experts — AWS ML Blog · 2026-10-06