Cheapest hardware to run Qwen3.8 27B at Q8? Ascend 310 out of stock, MI50 questioned
Snoo-2768 · reddit · 2026-09-02
A Reddit user asks for cost-effective hardware to run large models locally, targeting at least 20 tokens/sec on Qwen3.8 27B at Q8 quantization, with price as the top constraint.
- One candidate is the Huawei Ascend 310 series with 96GB memory, but it is completely out of stock everywhere
- Also considering the AMD Instinct MI50, unsure whether it is viable or garbage
- The post is primarily a question; discussion is in the comments
More from Infra
- Weaviate ships HFresh, a disk-based vector index inspired by SPFresh for large shifting collections — victorialslocum · 2026-09-02
- antirez: DeepSeek v4 Flash is still the king of local inference, now with vision — antirez · 2026-09-02
- Self-hosted LLM with SLERP-merged GRPO experts outperforms larger baseline, serving half of production traffic — t-tech · 2026-09-02
- AI Infrastructure Night event in San Francisco — glcst · 2026-09-02
- Google signs 396 MW geothermal deal to power AI amid energy crunch — VraserX · 2026-09-02
- Local Model Suitability MCP: Cuts Costs via Local Inference — modelcontextprotocol · 2026-09-02