Qwen 3 30B Hits 1000+ tok/s on M5 Max, Highlighting Apple's Edge in Local AI

StefanoGogioso · x · 2026-08-11

A researcher shared their experience running finetuned Qwen3-30B-A3B models locally on an Apple M5 Max laptop. Achieving over 1000 tokens/second in batched processing unlocks workflows that were unthinkable a year ago. The author praised Apple as the undisputed leader in consumer hardware for local intelligence, boasting a massive competitive moat.

Original post →

More from Infra

Infra channel →