Running Qwen 3.8 27B on M2 MacBook Pro 32GB: full tutorial and benchmarks
boutell · reddit · 2026-08-17
The author successfully runs Qwen 3.8 27B on an M2 MacBook Pro with 32GB RAM, sharing a step-by-step guide: build llama.cpp from source, download GGUF and vision adapter, close apps to free memory, and launch with specific flags. Benchmarks: 21.9 t/s prompt, 8.6 t/s generation. Includes aliases, API server setup, and integration with coding tools.
More from coding & agent
- Open-source tutorial: Build an AI telephony agent with VideoSDK and SIP trunking for inbound/outbound calls — tom_doerr · 2026-08-17
- Grok 4.6 tops VISTA benchmark, turning Figma designs into web apps at $2.38 per task — XFreeze · 2026-08-17
- From Loop to Graph: The Definitive Architecture Guide Beyond Single-Agent — iamrobotbear · 2026-08-17
- User uses Codex to automate scraping 5,302 X bookmarks dating back to 2014 — emollick · 2026-08-17
- Agent audit finds 9 bugs: build gates should list exemptions, not obligations — Federal-Teaching2800 · 2026-08-17
- Hermes Agent Builds Game Locally on 32GB GPU with Open-Weight 27B Model — Teknium · 2026-08-17