Muse Glimmer 30B Hits 25 tok/s In-Browser on M4 Max via Custom WebGPU Kernels

xenovatech · reddit · 2026-08-12

A developer successfully ran the Muse Glimmer 30B model locally inside a browser on an M4 Max Mac. By utilizing custom WebGPU kernels, the model achieved a generation speed of approximately 25 tokens per second. This performance is on par with running the model natively using llama.cpp, highlighting the significant potential of WebGPU for on-device LLM inference.

Related event: M4 Max Runs 30B Model at 25 tok/s in Browser via WebGPU(2 posts)→

Original post →

More from Infra

Infra channel →