MiniMax says MI355X is nearing B200 for serving its 428B multimodal model
salykova_ · x · 2026-07-23
MiniMax says it is working with hardware vendors to improve inference performance and keep the stack open. In a quoted note, Ryan Lee says AMD MI355X is now close to Nvidia B200 for serving MiniMax-M3.
The company points to two technical blogs covering full-stack ATOM + ATOMesh co-optimization for its 428B MSA sparse MoE multimodal base model. The optimization details include:
- EAGLE3 speculative decoding to reduce generation latency
- MSA-aligned Page-16 KV cache for 1M long context
- Layer-wise FP8 quantization to avoid precision loss in vision/MoE parts
- Distributed P/D separation for large clusters
The post is mainly about serving efficiency, long-context infrastructure, and hardware/software co-design.
More from Infra
- GPU racks are stalling on cold-plate and CDU capacity, not chip supply — tengyanAI · 2026-07-23
- Celeris launches a lab to build an LLM with microsecond response times — timshi_ai · 2026-07-23
- Reddit thread asks how to catch runaway agents before they blow the budget — Designer_Power3691 · 2026-07-23
- A curated guide to LLM cache management spans KV cache, batching, and decoding — gaganghotra_ · 2026-07-23
- RunPod users get a Chrome extension that notifies and auto-claims saved pods — Particular-Repair895 · 2026-07-23
- Aurora launches an open-source Go gateway for routing and securing LLM traffic — Select-Medicine-9310 · 2026-07-23