Running 104GB Qwen3.8-Flash-Next on 48GB Mac at ~12 tok/s

yogthos · reddit · 2026-09-02

A Reddit user shared their experience running the Qwen3 8B Flash Next model (quantized to 104GB) on a 48GB Mac, achieving an inference speed of approximately 12 tokens per second.

The implementation relies on the slotstream project on GitHub. This demonstrates the viability of running larger parameter models on memory-constrained devices via quantization techniques.

Original post →

More from Infra

Infra channel →