64GB Mac runs ~100GB Qwen model via oMLX with 130k context

chibop1 · reddit · 2026-10-11

A Reddit user reports running a large quantized Qwen model (Qwen3.8-Flash-Next-oQ4e-mtp) on an M3 Max with 64GB RAM using the latest oMLX commit — something that failed two weeks ago. About 58GB is allocated to the GPU and it handles up to 130k context.

Test session stats: 13 requests, 321k total prefill tokens, 90.9% cache efficiency, 140 tok/s prompt processing (excluding cached) and 15.8 tok/s generation.

Key oMLX settings: Aggressive memory guard, 2GB hot cache limit, Lightning MTP on, MoE Expert Offload with 50% resident experts.

Original post →

More from Infra

Infra channel →