Running 30B Model in Chrome on Mac Studio Hits 32 tok/s

MaziyarPanahi · x · 2026-08-12

A developer tested loading a 30B dense model (Muse Glimmer) entirely within Chrome on a Mac Studio. While running a clinical AI pipeline—including PII detection, clinical NER, and relation extraction—the inference speed reached 32 tokens per second. This real-world test highlights the massive leaps in local and in-browser AI deployment. The author noted that a similar demo for Qwen3.8 27B will follow later today.

Related event: Muse Glimmer 30B Tested: A New Standard for Local Multimodal Agents(14 posts)→

Original post →

More from Infra

Infra channel →