Running 30B LLM in Chrome at 32 tok/s: Local AI Takes a Massive Leap

MaziyarPanahi · x · 2026-08-12

A developer tested loading a quantized 30B dense model entirely within the Chrome browser on a Mac Studio, achieving an impressive 32 tokens/second. The model successfully ran a complex clinical AI pipeline—including PII detection and clinical NER—locally, demonstrating a massive leap in local AI inference capabilities.

Related event: Muse Glimmer 30B Tested: A New Standard for Local Multimodal Agents(14 posts)→

Original post →

More from Infra

Infra channel →