Can Qwen Flash Next Run on a 64GB RAM iGPU Mini PC? One User's Experiment
SomeITGuyLA · reddit · 2026-10-03
A Reddit user explores running Qwen Flash Next on a non-Mac mini PC with 64GB unified RAM and a 780M iGPU, noting others have run it with 12GB VRAM + 64GB RAM and on Macs. He currently runs 125B Ling 3.0 Flash at Q2 quants via llama.cpp (vulkan), and proposes offloading ngrams to SSD — unsupported in llama.cpp, and other engines lack vulkan support.
More from Infra
- Modal VM Sandboxes hit GA: demo runs Docker Compose apps, tests and coding agents — charles_irl · 2026-10-03
- Community poll of 793 MLX users: oMLX wins at 55.5%, dominating Ultra chips — HankYeomans · 2026-10-03
- Local AI user asks: why does everyone enjoy taming the beast of complexity? — kathi7 · 2026-10-03
- AI gateways emerge as key control layer, with eight core capabilities to govern production AI stacks — goyalshaliniuk · 2026-10-03
- 8GB VRAM folks' daily prayer for Qwen4 35B A3B, settling for Gemma 26B QAT — RobustLokiX · 2026-10-03
- Two 300B MoE models on one 128GB Strix Halo: Kyojin engine hits 44 tok/s decode — Yaniss916 · 2026-10-03