MacBook M5 Pro 48GB runs Qwen Flash at 13-14 tok/s — what's your local setup?
carloslfu · reddit · 2026-09-16
A user kicks off a discussion on the best local models for Mac: on a MacBook M5 Pro with 48GB, they run Qwen 3.8 Flash at 13-14 tok/s using 21-24GB of memory.
The thread invites other Mac owners to share their specs and preferred local models, offering a rough comparison across hardware tiers.
More from Infra
- Lithos: stop treating AI benchmarks as proof, define your own metrics — JiaZhihao · 2026-09-16
- Cheap third-party open LLM providers ranked, coral bricks tops the list — opensourcecolumbus · 2026-09-16
- Could an Xbox Series X cluster with 160GB unified memory host local LLMs? A Reddit thought experiment — RecursiveCTE · 2026-09-16
- Rob Mulla marks one year at Google scaling TPU inference with vLLM — Rob_Mulla · 2026-09-16
- Jev claims new frontier model 40-400x cheaper and 20-200x faster — albelfio · 2026-09-16
- Nvidia B200 Compute Prices Jump 21% in a Month as AI Demand Outstrips Supply — JOBhakdi · 2026-09-16