Running Qwen LLMs on 64GB RAM: Tested
misha1350 · reddit · 2026-07-16
The author shares a hands-on comparison of running large models locally on a 64GB RAM machine, focusing on two quantized models: Qwen3 Next 80B and Qwen3.5 122B A10B.
Key takeaways:
- The 80B model runs at roughly 8.5 tok/s on 64GB RAM, offering faster generation speeds.
- The 122B A10B can also fit into the same memory footprint via more aggressive quantization, but its speed drops to about 2.9 tok/s, though it delivers better response quality and internal knowledge.
- The author notes that while larger sparse/dense models are more reliable for niche questions, their prompt processing speeds are too slow for agentic workflows.
- They also point out a lack of a "sweet spot" model for the 32–48GB memory tier; a well-suited mid-sized model would round out the local LLM experience.
- Finally, the REAP version of Qwen3.5 122B A10B performed even worse, failing to resolve the core speed bottleneck.
More from Infra
- NeurIPS 2026 workshop will focus on on-device intelligence and local execution — YiMaTweets · 2026-07-21
- How to build a PostgreSQL-backed semantic search pipeline with pgvector and Ollama — KhuyenTran16 · 2026-07-21
- NeurIPS 2026 workshop calls papers on on-device intelligence — YiMaTweets · 2026-07-21
- Milled from Solid Aluminum: AI Rig Multi-GPU Case for Local Compute — dee_hw · 2026-07-21
- FutureCaribbean’s Buildathon offers $50K, H200 compute, and an NYSE pitch — HeyAmit_ · 2026-07-21
- A new series tests which data-science workflows can run on GPUs today — pandeyparul · 2026-07-21