Running 104GB Qwen3.8-Flash-Next on 48GB Mac at ~12 tok/s
yogthos · reddit · 2026-09-02
A Reddit user shared their experience running the Qwen3 8B Flash Next model (quantized to 104GB) on a 48GB Mac, achieving an inference speed of approximately 12 tokens per second.
The implementation relies on the slotstream project on GitHub. This demonstrates the viability of running larger parameter models on memory-constrained devices via quantization techniques.
More from Infra
- Ilya warns: Neoclouds must strengthen security against AI takeovers — _AndrewZhao · 2026-09-02
- User questions reliance on AI vendor lacking sandbox expertise — basedjensen · 2026-09-02
- Deep Dive into OpenAI's Jalapeño Chip Architecture: A System-Level Analysis — thehiphopswami · 2026-09-02
- Nous Research Launches Portal to Unify Agent Ecosystem — Teknium · 2026-09-02
- Google signs 396 MW geothermal deal to power potential Utah data center — natesiggard · 2026-09-02
- YC Paper Club Call: Optical Compute, Diamond Chips, and Bio-GPUs — ycombinator · 2026-09-02