Maia 200 hits 99.69% of BF16 matmul peak by programming data movement

Maia 200: A Software Defined Dataflow System for Large-scale AI Acceleration

Sherry Xu, Marco Heddes, Jackson Peng, Tom Savell, Monica Tang, Prashant Ranjan, Jesse Benson, Ofer Dekel, Saurabh Dighe, Anupama Kurpad, Artour Levin, Matthew Mattina, George Petre, Cheng Tang, Yuan Yu, Li Zhang, Torsten Hoefler

cs.AR, cs.AI, cs.DC, cs.ET, cs.LG

2026-08-25

Microsoft's Maia 200 SDLA chip posts 10,145 Tflop/s FP4 at 750W, 99.69% of BF16 matmul peak, and 2,434 tokens/s on Qwen 2.5 7B decode.

What problem this solves

Inference now burns most of the accelerator cycles on the planet. Microsoft's conservative floor is 1.2 trillion tokens per day. An 8B model at that volume is about 6.85 exaflop/s sustained; a 400B model is 0.3 zettaflop/s. The scarce resource is not another tensor core. It is moving data, then computing on it immediately. GPU SIMT still treats threads as the first-class object. DMA and warp scheduling live in firmware. Expert programmers fake a dataflow machine with warp specialization, splitting loads and math across different warps.

Maia 200 is Microsoft's second-generation in-house accelerator, already in production on Azure. The paper casts it as an abstract machine: a Software Defined Locally Accessed Dataflow Architecture (SDLA). Control and data paths are programmed separately. On-chip memories sit next to the units that use them. The claim to test is whether that contract, on real silicon, can run close to peak.

Method

SDLA borrows Flynn's two-axis habit. One axis is where data lives (local vs global). The other is who issues movement (load/store vs a separate software-defined stream). Shared-memory CPUs sit in LSGA; GPUs and Cerebras in LSLA; CPUs with programmable DMA in SDGA. Maia sits in SDLA: local addressing plus an explicit data-movement instruction stream. The contract has three rules. Expose parallel engines for movement, conversion, and sync. Let software manage distributed, specialized memories. Keep performance deterministic enough to schedule.

The chip is TSMC 3 nm, more than 140 billion transistors, a near-reticle 26×33 mm die in a 75×75 mm CoWoS-S package, 750 W SoC TDP, six HBM3e stacks at 7 TiB/s. Four clusters, nine or ten tiles each. A tile pairs a Tile Tensor Unit (TTU) with a Tile Vector Processor and 3 MiB of private SRAM. The TTU does 65,536 FP4 MACs per cycle; at 2 GHz that is about 262 Tflop/s per unit. C/C++ control programs run on a three-level processor hierarchy (device, cluster, tile). The data path speaks a Dataflow ISA: a macro-instruction waits on up to two semaphores and signals up to two when it finishes. SRAM takes under 20% of the die. Caches, the authors note, spend 30-35% of their area on tags; most AI kernels are data-oblivious, so the compiler can place every byte.

The network is 28 integrated 400 Gbps Ethernet ANCs, 1.4 TB/s full duplex, running Microsoft's ATLv2, which later fed Ultra Ethernet. 6,144 chips is the current inference design point: 62 exaflop/s FP4, 43 PiB of memory, 8.6 PiB/s of Ethernet.

Results

Peak is 10,145 Tflop/s FP4 and 5,072 Tflop/s FP8 inside 750 W, or 13.3 and 6.7 Tflop/W. Internal accounting claims 30% lower TCO and 15% less energy than any other AI accelerator in Microsoft's fleet. The paper refuses a public GPU bake-off.

With nine tiles per cluster at 2 GHz unthrottled, BF16 peak is 1,180 Tflop/s and FP8 is 4,785 Tflop/s. Across 6,143 inference-relevant GEMMs, loaded from HBM and written back:

SetupMetricResult
BF16, compute-boundvs peakup to 99.69%
BF16, >58 Tflop of workvs peak>90%
BF16, memory-boundvs peak bandwidthup to 51.4%
FP8, compute-boundvs peakup to 96%
Allgather, 8 chipsvs latency / bandwidth SoL78% / 94%

End to end, Qwen 2.5 7B decode of token 16,385 after 16,384 tokens of context, KV cache 939.52 MiB. A plain PyTorch path with stock kernels and no fusion hits 2,434 tokens/s, more than 70% of the authors' estimated ceiling.

Why it matters

For anyone who has rewritten FlashAttention or chased TMA across GPU generations, this paper promotes hand-scheduled data movement from a trick to a machine contract. Hopper's TMA still has to dance with warps; Blackwell replaced WGMMA with UMMA and broke compatibility to make that dance less painful. Maia treated local scratchpads and async DMA as first-class from day one, and the compute-bound GEMM numbers show it. The cost is a spatial programming model: PyTorch and Triton on top, C++ and semaphores in the hot kernels.

The TCO and energy claims are internal. Outsiders cannot audit them. What can be audited is the roofline the authors drew themselves: near-full utilization when the math is the bottleneck, about half the peak bandwidth when it is not.

Limitations

The software stack is almost undescribed, so how the compiler fills scratchpads is a black box. Matmul and Allgather are microbenchmarks. The only end-to-end number is 7B decode: no prefill, no MoE, no multi-chip generation throughput. The authors say they skipped a direct GPU comparison on purpose. The 30% TCO and 15% energy figures name neither the comparison chip nor the utilization and electricity assumptions. 6,144 chips is a design point, not a measured system in this paper. For dynamic batching, sparsity, and agent workflows, the static-planning advantage of SDLA shrinks. MoE routing is described as a few-cycle hop inside a tile, with no speed number attached.

Terms

Source

What people are saying

Related papers

All paper explainers