wllama V3 ships WebGPU, multimodal and tool calling for in-browser llama.cpp inference
ngxson · x · 2026-09-19
wllama, ngxson's WebAssembly binding for llama.cpp (1.2k GitHub stars), just released V3 with three headline features: WebGPU support, multimodal input (image and audio), and tool calling. It exposes a fully-typed OpenAI-compatible API, and the author built the openjev demo on top of it.
Weights load from Hugging Face and cache in the browser with inputs never leaving the page, making it a solid base for backend-free local AI web apps.
Related event: wllama V3 and openjev bring local LLMs to the browser(3 posts)→
More from Infra
- TIL: AMD compute profiler uses violet vs purple to mark divergent branches in waves — salykova_ · 2026-09-19
- CDNA kernel devs say assembly with syntactic sugar may beat the compiler — salykova_ · 2026-09-19
- Analyst sees 20-30% gap between AI accelerator demand and energized capacity — BenBajarin · 2026-09-19
- PyTorch Conference to feature Muon optimizer, TorchJD and LLM serving tools from Meta, Google, NVIDIA — PyTorch · 2026-09-19
- AMD MI355X sustains 2x the concurrent agents of MI300X in Signal65 PINNACLE tests, up to 6.5x on MiniMax-M3 — ryanshrout · 2026-09-19
- ASICs to outship GPUs next year, but custom silicon remains a hyperscaler oligopoly — ai · 2026-09-19