Two 96GB Huawei Ascend cards run Qwen locally: from 1 tok/s to 30 tok/s

matteiuspi · reddit · 2026-10-02

The author built a local inference machine around two Huawei Atlas 300I Duo cards (4× Ascend 310P3 devices, 192GB LPDDR4X total, 172GiB visible) running Qwen3.8 Flash-Next.

Hardware notes

Software work

Original post →

More from Infra

Infra channel →