Running 30B Model in Chrome Hits 32 tok/s on Mac Studio, Showcasing Local AI Power
MaziyarPanahi · x · 2026-08-12
A developer achieved 32 tokens/s running the 30B dense model Muse Glimmer entirely inside the Chrome browser on a Mac Studio.
The local setup successfully powered a clinical AI pipeline featuring PII detection, clinical NER, and relation extraction. This mind-blowing performance highlights the massive strides made in local AI inference technology.
Related event: Meta's Open-Source Muse Glimmer 30B Impresses in On-Device Agent Tests(14 posts)→
More from Infra
- Data centers leave little water for residents — CtrlAltDwayne · 2026-08-26
- OpenAI reveals custom inference chip Jalapeño with higher throughput and lower latency — Moh1tAgarwal · 2026-08-26
- Mixedbread on retrieval scaling laws: co-designing models and vector DBs — lateinteraction · 2026-08-26
- Data Center Backlash Not Driven by Anti-Tech Sentiment — AndyMasley · 2026-08-26
- AI Agent Security Market: Can Zscaler Become the Default Control Plane? — thedealdirector · 2026-08-26
- Running Qwen 27B on RTX 3060+2060 Yields Only 5-6 TPS — sheriffoftiltover · 2026-08-26