Running LLMs Fully Client-Side: llama.cpp Compiled to WASM with WebGPU Proof of Concept

Numerous-Fan8138 · reddit · 2026-10-06

The author shows a weekend experiment: a proof-of-concept running an LLM entirely client-side in the browser, by compiling llama.cpp to WASM with WebGPU. A demo video is included. All inference happens on the user's device with no server involved — a hands-on exploration of client-side/local LLM deployment.

Related event: llama.cpp runs LLMs fully in-browser via WASM and WebGPU(3 posts)→

Original post →

More from Infra

Infra channel →