Microsoft's Maia 200 Paper Released: Software-Defined Dataflow for High-Performance Inference
thoefler · x · 2026-08-26
Microsoft presented the Maia 200 accelerator at HotChips 26, detailing how its Software-Defined Dataflow architecture delivers peak inference performance for production LLMs, accompanied by a full paper release. Maia 200 represents a new class of Software Defined Locally Accessed Dataflow Architectures (SDLA), explicitly programming dataflow engines to orchestrate specialized memories and data movement engines. This approach shifts focus from thread-centric to data-movement-centric architectures. The chip delivers 10,145 TFLOPS (FP4) and 5,072 TFLOPS (FP8) within a 750W TDP and 7 TB/s HBM bandwidth, targeting significant cost and energy savings for large-scale AI inference workloads.
More from Infra
- Linux Foundation Celebrates 35th Anniversary, Open Source Targets AI — 0xsachi · 2026-08-26
- Open Source Facefusion Android App Runs Video Face Swap on Snapdragon NPU — Few_Caregiver8134 · 2026-08-26
- Apple M5 Mac Studio page features LM Studio for local AI — mattturck · 2026-08-26
- Chinese chip packaging firms invest $2.2B in expansion fueled by AI boom — pstAsiatech · 2026-08-26
- Can I run MiniMax H3 locally on an RTX 2060 with 6GB VRAM? — Amjad_K · 2026-08-26
- a16z Partner: Distinction between training and inference will become meaningless — Kyrannio · 2026-08-26