64GB Mac runs ~100GB Qwen model via oMLX with 130k context
chibop1 · reddit · 2026-10-11
A Reddit user reports running a large quantized Qwen model (Qwen3.8-Flash-Next-oQ4e-mtp) on an M3 Max with 64GB RAM using the latest oMLX commit — something that failed two weeks ago. About 58GB is allocated to the GPU and it handles up to 130k context.
Test session stats: 13 requests, 321k total prefill tokens, 90.9% cache efficiency, 140 tok/s prompt processing (excluding cached) and 15.8 tok/s generation.
Key oMLX settings: Aggressive memory guard, 2GB hot cache limit, Lightning MTP on, MoE Expert Offload with 50% resident experts.
More from Infra
- Running 426GB MXFP4 DeepSeek on 192GB VRAM to power a 4-sub-agent code review — HankYeomans · 2026-10-11
- awesome-local-ai: a task-sorted local AI tool list built for coding agents — blaizedsouza · 2026-10-11
- A single prompt made an AI find a bug eating 84% of backend CPU — zealcaiden · 2026-10-11
- ComfyUI on a $899 M4 Mac Mini 16GB: Full Benchmark Results and Workflows — FaatmanSlim · 2026-10-11
- Asana cut a Codex agent's cost 4x to $0.47 per run with an 89% prompt cache hit rate — daniel_mac8 · 2026-10-11
- 'The most important chart' of the AI economy: model layer commoditizing, compute demand soaring — dnlkwk · 2026-10-11