Running 30B Models Locally in Chrome at Over 30 tok/s

MaziyarPanahi · x · 2026-08-13

A developer successfully ran the Muse Glimmer 30B dense model entirely locally inside the Chrome browser. It achieved an inference speed of 31.6 tokens/second while handling tasks like clinical PII, Named Entity Recognition (NER), and relation extraction. This demonstrates the significant potential for deploying large AI models directly in browsers.

Original post →

More from Infra

Infra channel →