Dev gets llama.cpp compiling with WebGPU, runs an LLM fully in the browser
Numerous-Fan8138 · reddit · 2026-10-06
A Redditor spent the weekend getting llama.cpp to compile with the WebGPU backend and published a working POC: an LLM running entirely inside the browser with no server.
- Current state: functional proof-of-concept repo, in-browser inference
- Next step: turn the LLM into an agent that can interact with the website or web app it lives in
- The author is soliciting feedback on whether the direction makes sense
A useful starting point for anyone exploring serverless, fully client-side local LLM/agent setups.
Related event: llama.cpp runs LLMs fully in-browser via WASM and WebGPU(3 posts)→
More from coding & agent
- AgentHopper: a cross-agent 'AI virus' built on chained prompt injection — wunderwuzzi23 · 2026-10-06
- How often should you run evals? Hamel Husain lays out a 3-factor tradeoff framework — HamelHusain · 2026-10-06
- Gemini 3 Flash burned $14 and failed a $1 agent task Pro finished in 85 steps — XIFAQ · 2026-10-06
- This builder spends 5x more on scraping APIs than AI models — data is the edge — EXM7777 · 2026-10-06
- Interfaze open-weights MoA model for OCR, speech and GUI grounding on one 80GB GPU — charles_irl · 2026-10-06
- 10-person DTC brand runs on 16 AI agents: 90 improvements in 15 weeks, now sells AI setup to other brands — jacob_posel · 2026-10-06