AMD MI455X Architecture Breakdown: First Rack-Native GPU with HBM4
ryanshrout · x · 2026-07-24
AMD has introduced the Instinct MI455X GPU, designed specifically for rack-scale AI rather than competing purely at the chip level. The architecture features several major breakthroughs:
- Compute & Bandwidth: Delivers roughly 40 PFLOPS of MXFP4 compute. It increases bandwidth to 3.6 TB/s via 36 UALoE links, allowing 72 GPUs within a rack to act as a single shared-memory system.
- Memory Advantage: Packs 432 GB of HBM4 with 23.3 TB/s bandwidth and grows the L2 cache to 192 MB, heavily optimizing bottlenecks for Mixture-of-Experts models and long-context inference.
- Architecture Overhaul: Transitions from Wave64 to a native Wave32 execution model, increases addressable registers to 1,024 per thread, and introduces a dedicated Tensor Data Mover.
- Advanced Packaging: Utilizes TSMC's 2nm process for 8 compute dies, combined with 3D stacking and CoWoS-L packaging, totaling 320 billion transistors.
More from Infra
- Leaked DeepSeek transcript says the company has only 20,000 H-equivalent cards — fiiiiiist · 2026-07-24
- Investor Critique: Etching Transformers Into Silicon Is Inherently Limiting — JosephJacks_ · 2026-07-24
- Huawei’s 4:1 GB300 claim shrinks to 2:1 on memory bandwidth, thread says — zephyr_z9 · 2026-07-24
- NVIDIA introduces NVFP4 for faster LLM inference with less GPU memory — NVIDIA Developer · 2026-07-24
- DeepSeek-V4-Flash reaches 105 tok/s on two 4090D cards after Triton kernel rewrites — iSevenDays · 2026-07-24
- Analyst: We Are Still in the First Generation of Rack-Scale AI Compute — BenBajarin · 2026-07-24