Qwen3.8-9B MLX port runs on 16GB Macs with fast speed
alexcovo_eth · x · 2026-08-18
The 4-bit MLX version of Qwen3.8-9B Abliterated is now available, optimized for Apple Silicon. Benchmarks on M5 Max show 1028 tok/s prefill and 44.6 tok/s generation with 13.05 GB peak memory, making it runnable on 16GB Macs. It is a third-party distillation of Qwen3.5-9B with safety refusals removed.
More from Infra
- Model Prices Collapse, but Gateway Fees Remain High — AccBalanced · 2026-08-18
- Performance Analysis of Mixing RTX 3080 and A2000 for Stable Diffusion — djekler · 2026-08-18
- Snapdragon X2 Elite Extreme Doubles AI Performance Over Rivals — ryanshrout · 2026-08-18
- Salvaging parts: building a local AI rig with 64GB RAM and a €1700 budget — joquinjack · 2026-08-18
- Mesh LLM: A Third Option Between Crypto Rigs and Cloud Subscriptions — alex_verem · 2026-08-18
- Overclocking VRAM on 4x RTX 5060Ti for LLM inference — Ok-Breakfast1878 · 2026-08-18