Running 30B Model in Chrome on Mac Studio Hits 32 tok/s
MaziyarPanahi · x · 2026-08-12
A developer tested loading a 30B dense model (Muse Glimmer) entirely within Chrome on a Mac Studio. While running a clinical AI pipeline—including PII detection, clinical NER, and relation extraction—the inference speed reached 32 tokens per second. This real-world test highlights the massive leaps in local and in-browser AI deployment. The author noted that a similar demo for Qwen3.8 27B will follow later today.
Related event: Muse Glimmer 30B Tested: A New Standard for Local Multimodal Agents(14 posts)→
More from Infra
- AI Compute Boom Drives TL20 Tech Stocks Up 59% Year-to-Date — TiernanRayTech · 2026-08-12
- How Vercel Migrated Its Core Database Handling 6,000 Deployments Per Minute — evilrabbit_ · 2026-08-12
- Analyst: Server CPU Market Bracing for Unprecedented S-Curve Leap — BenBajarin · 2026-08-12
- ComfyUI Workflows Crawl on 128GB DGX Spark vs RTX 4090 — jungseungoh97 · 2026-08-12
- Server CPU Market to Reach $220B by 2030, Potentially Split in Thirds — BenBajarin · 2026-08-12
- HuggingFace Models See ~30% Sustained Drop in Daily Downloads, Likely Due to Filtering Changes — natolambert · 2026-08-12