Muse Glimmer 30B Runs in Browser at Native llama.cpp Speed

nicodotdev · x · 2026-08-12

Meta's new Muse Glimmer 30B model can now run locally in the browser via WebGPU. Tests show it achieves about 25 tok/s on an M4 Max, matching native llama.cpp speeds without requiring any installation.

Related event: M4 Max Runs 30B Model at 25 tok/s in Browser via WebGPU(2 posts)→

Original post →

More from Models

Models channel →