AMD MI100 User Report: Cooling Struggles and Subpar Inference Speeds
faisalkl · reddit · 2026-09-02
A user reported a poor experience using the AMD Instinct MI100 32GB for local LLM inference. Key issues include extreme cooling difficulties, requiring high-static-pressure fans to avoid thermal throttling at 175-200W. Performance-wise, despite 1.2TB/s theoretical bandwidth, real-world token generation (20-40 t/s) with Qwen models lags behind expectations and even lower-bandwidth consumer GPUs like the R9 7000 series. The user suspects core rate limiting or software optimization issues and seeks advice from the community.
More from Infra
- Zyphra open-sources PUFFER, a CPU-based deduplication system 35x faster — bclavie · 2026-09-02
- Dell reports record $47B quarterly revenue, raises outlook on AI infrastructure demand — Polymarket · 2026-09-02
- Nvidia cuts Rubin Ultra memory to 192GB as HBM costs soar to 40% of TCO — rwang07 · 2026-09-02
- SpaceX data center team shakeup: Musk replaces leaders with rocket, satellite internet execs — kyliebytes · 2026-09-02
- Seeking Datacenter-Grade OCS — jwt0625 · 2026-09-02
- GB10 price hike sparks debate: Is Mac Studio the best value for compute? — geekender · 2026-09-02