MiniMax says AMD MI355X is now close to Nvidia B200 in model serving
hongyangzh · x · 2026-07-24
MiniMax says AMD MI355X is nearing Nvidia B200 in serving
A repost from MiniMax claims AMD MI355X is now close to Nvidia B200 for serving MiniMax-M3.
The company says it published two technical blog posts describing a full-stack co-optimization of its ATOM + ATOMesh serving system for a 428B sparse MoE multimodal base model.
The post highlights several engineering changes:
- EAGLE3 speculative decoding to reduce generation latency
- MSA-aligned Page-16 KV cache for 1M ultra-long context
- Layer-wise FP8 quantization to avoid precision loss in vision/MoE components
- Distributed P/D separation for large-scale clusters
The main takeaway is that AMD hardware is getting closer to Nvidia performance in a real large-scale serving setup.
Related event: MiniMax Claims AMD MI355X Serving Performance Nears Nvidia B200(2 posts)→
More from Infra
- Alphabet revenue rose 24% to $119.8B as Google Cloud grew 82% — minsuk_chang · 2026-07-24
- Taylor Lorenz says the AI data-center backlash has lost nuance — AndyMasley · 2026-07-24
- Vercel says Python functions now start 2x faster after build-time bytecode precompilation — cramforce · 2026-07-24
- Google’s new servers may be feeding Search, YouTube and ads—not just Cloud — aronchick · 2026-07-24
- Advanced packaging, not just leading nodes, is the AI bottleneck — Beth_Kindig · 2026-07-24
- An agent can now pay for its own data using the x402 protocol — kleffew94 · 2026-07-24