Microsoft's Maia 200 Paper Released: Software-Defined Dataflow for High-Performance Inference

thoefler · x · 2026-08-26

Microsoft presented the Maia 200 accelerator at HotChips 26, detailing how its Software-Defined Dataflow architecture delivers peak inference performance for production LLMs, accompanied by a full paper release. Maia 200 represents a new class of Software Defined Locally Accessed Dataflow Architectures (SDLA), explicitly programming dataflow engines to orchestrate specialized memories and data movement engines. This approach shifts focus from thread-centric to data-movement-centric architectures. The chip delivers 10,145 TFLOPS (FP4) and 5,072 TFLOPS (FP8) within a 750W TDP and 7 TB/s HBM bandwidth, targeting significant cost and energy savings for large-scale AI inference workloads.

Original post →

More from Infra

Infra channel →