Qwen3.8-9B MLX port runs on 16GB Macs with fast speed

alexcovo_eth · x · 2026-08-18

The 4-bit MLX version of Qwen3.8-9B Abliterated is now available, optimized for Apple Silicon. Benchmarks on M5 Max show 1028 tok/s prefill and 44.6 tok/s generation with 13.05 GB peak memory, making it runnable on 16GB Macs. It is a third-party distillation of Qwen3.5-9B with safety refusals removed.

Original post →

More from Infra

Infra channel →