Two 96GB Huawei Ascend cards run Qwen locally: from 1 tok/s to 30 tok/s
matteiuspi · reddit · 2026-10-02
The author built a local inference machine around two Huawei Atlas 300I Duo cards (4× Ascend 310P3 devices, 192GB LPDDR4X total, 172GiB visible) running Qwen3.8 Flash-Next.
Hardware notes
- Each card is really two chips with 48GB local each; tensor/expert parallelism required — think four 48GB ranks, not two 96GB GPUs
- Memory is LPDDR4X, not HBM; 408 GB/s bandwidth, PCIe Gen4 x16, single-slot 150W passive design
- Passive cooling works with directed airflow and AC cooling: 72–78°C under multi-hour loads
Software work
- Started incoherent at 1 tok/s; two weeks of vLLM/vLLM Ascend work on drivers, model architecture support, memory layout, custom operators, and async state transitions
- Now 30 tok/s single request, 61 tok/s aggregate at 4-way concurrency; completed the full 198-question GPQA Diamond set
- Planned: 3 cards, 288GB nameplate, 450W total board power
More from Infra
- AI Capex Now Exceeds the Railroad Boom's GDP Share, Yet Demand Lags — FinanceYF5 · 2026-10-02
- Real-time AI serving costs up to 56x more than needed — one engine hits 56 sessions per H100 — Ok_boss_labrunz · 2026-10-02
- Cloudflare launches SQL API to query Workers logs and traces, replacing GraphQL plans — irvinebroque · 2026-10-02
- Cloudflare launches Web Search API via AI Gateway with Exa, Linkup and Ceramic — michellechen · 2026-10-02
- Lightpanda 1.0 ships: a Zig-built browser for AI agents, out of beta — jedisct1 · 2026-10-02
- Qwen3.8-27B coder quant fits a 24GB GPU with 262k context at 40 t/s — W61k3r · 2026-10-02