AMD MI455X Architecture Breakdown: First Rack-Native GPU with HBM4
ryanshrout · x · 2026-07-24
AMD has introduced the Instinct MI455X GPU, designed specifically for rack-scale AI rather than competing purely at the chip level. The architecture features several major breakthroughs:
- Compute & Bandwidth: Delivers roughly 40 PFLOPS of MXFP4 compute. It increases bandwidth to 3.6 TB/s via 36 UALoE links, allowing 72 GPUs within a rack to act as a single shared-memory system.
- Memory Advantage: Packs 432 GB of HBM4 with 23.3 TB/s bandwidth and grows the L2 cache to 192 MB, heavily optimizing bottlenecks for Mixture-of-Experts models and long-context inference.
- Architecture Overhaul: Transitions from Wave64 to a native Wave32 execution model, increases addressable registers to 1,024 per thread, and introduces a dedicated Tensor Data Mover.
- Advanced Packaging: Utilizes TSMC's 2nm process for 8 compute dies, combined with 3D stacking and CoWoS-L packaging, totaling 320 billion transistors.
Related event: AMD Launches MI455X and Helios, Escalating AI Compute to Rack-Scale(9 posts)→
More from Infra
- LLM Serving Metrics Thread: Why TPOT and Uptime Make or Break User Experience — abhijithneil · 2026-09-11
- PlanetScale launches sharded Postgres: 768 servers acting as one, 1PB scale — dhruv2038 · 2026-09-11
- Can a 7900 XTX 24GB run Qwen locally? Reddit seeks ROCm tok/s benchmarks — thenomadexplorerlife · 2026-09-11
- RTK Terminal Compression Cuts Tokens but Leaves Your AI Coding Bill Unchanged — Bartaseth · 2026-09-11
- SF Compute founder: buying compute is 'an absolutely awful experience' right now — IgorCarron · 2026-09-11
- SmolVM open-sources persistent computer infrastructure for agents that outlive chat sessions — aniketmaurya · 2026-09-11