Running 30B Models Locally in Chrome at Over 30 tok/s
MaziyarPanahi · x · 2026-08-13
A developer successfully ran the Muse Glimmer 30B dense model entirely locally inside the Chrome browser. It achieved an inference speed of 31.6 tokens/second while handling tasks like clinical PII, Named Entity Recognition (NER), and relation extraction. This demonstrates the significant potential for deploying large AI models directly in browsers.
More from Infra
- L&T and Together AI to Build 10,000-GPU NVIDIA B300 AI Factory in India — RoboBalaji · 2026-08-13
- Budget 1.5k EUR for local LLM hardware: Reddit user seeks advice on GPU choices — DisLLMs · 2026-08-13
- Volta, 7-Month-Old AI Infrastructure Startup, Raises $300M and Signs $10B Compute Deal with Anthropic — 创业邦 · 2026-08-13
- Lenovo Q1: AI Revenue Exceeds 63B Yuan with 360B+ AI Server Pipeline — 智东西 · 2026-08-13
- 23.5 TB VRAM and 288 GPUs: Is This Still 'Local AI'? — MaziyarPanahi · 2026-08-13
- Benchmark Reveals MCP Costs Up to 3x More Compute Than Plain Models — KitchenAmoeba4438 · 2026-08-13