AMD MI100 User Report: Cooling Struggles and Subpar Inference Speeds

faisalkl · reddit · 2026-09-02

A user reported a poor experience using the AMD Instinct MI100 32GB for local LLM inference. Key issues include extreme cooling difficulties, requiring high-static-pressure fans to avoid thermal throttling at 175-200W. Performance-wise, despite 1.2TB/s theoretical bandwidth, real-world token generation (20-40 t/s) with Qwen models lags behind expectations and even lower-bandwidth consumer GPUs like the R9 7000 series. The user suspects core rate limiting or software optimization issues and seeks advice from the community.

Original post →

More from Infra

Infra channel →