Running Qwen 3.8 27B on M2 MacBook Pro 32GB: full tutorial and benchmarks

boutell · reddit · 2026-08-17

The author successfully runs Qwen 3.8 27B on an M2 MacBook Pro with 32GB RAM, sharing a step-by-step guide: build llama.cpp from source, download GGUF and vision adapter, close apps to free memory, and launch with specific flags. Benchmarks: 21.9 t/s prompt, 8.6 t/s generation. Includes aliases, API server setup, and integration with coding tools.

Original post →

More from coding & agent

coding & agent channel →