Muse Glimmer 30B Hits 25 tok/s In-Browser on M4 Max via Custom WebGPU Kernels
xenovatech · reddit · 2026-08-12
A developer successfully ran the Muse Glimmer 30B model locally inside a browser on an M4 Max Mac. By utilizing custom WebGPU kernels, the model achieved a generation speed of approximately 25 tokens per second. This performance is on par with running the model natively using llama.cpp, highlighting the significant potential of WebGPU for on-device LLM inference.
Related event: M4 Max Runs 30B Model at 25 tok/s in Browser via WebGPU(2 posts)→
More from Infra
- $500B AI Infrastructure Funds May Shift to Neoclouds Over Hyperscalers — abhiadesai · 2026-08-12
- Mojo 1.0 Released: The Systems Language for the AI Era — clattner_llvm · 2026-08-12
- Nvidia's Switchyard Router Reshuffles AI Models Mid-Task, Cutting Costs to 1/3 — CackleRooster · 2026-08-12
- Data Center Tax Boom Leads to 10 Years of Property Tax Cuts in Virginia — robleclerc · 2026-08-12
- Breaking VM Barriers: Apple Silicon LLM Inference Runs 16x Faster — petrusenko_max · 2026-08-12
- Ling-3.0-flash Quantization Benchmarks: MoE Architecture Preserves Decode Speed — AcanthisittaOk1699 · 2026-08-12