Qwen 3 30B Hits 1000+ tok/s on M5 Max, Highlighting Apple's Edge in Local AI
StefanoGogioso · x · 2026-08-11
A researcher shared their experience running finetuned Qwen3-30B-A3B models locally on an Apple M5 Max laptop. Achieving over 1000 tokens/second in batched processing unlocks workflows that were unthinkable a year ago. The author praised Apple as the undisputed leader in consumer hardware for local intelligence, boasting a massive competitive moat.
More from Infra
- Demand for AI Gateways and Model Routing Sees Explosive Growth — shensi · 2026-08-11
- OpenAI's Letter to Texas Governor on Responsible AI Infrastructure — borowcy · 2026-08-11
- Testing DGX Spark with H3: First Text-to-Video Run via ComfyUI — allaboutai-kris · 2026-08-11
- China's First Domestic Front-end Lithography Tool Enters Production, Competing with 90s ASML — pstAsiatech · 2026-08-11
- Under 10% of Enterprises Scale AI; Compute Shortage to Persist — BenBajarin · 2026-08-11
- Amazon Backs Texas Gas Plant That May Become Top US Climate Polluter for AI — Ars Technica AI · 2026-08-11