Microsoft shows full-stack engineering around Maia 200 and Cobalt chips for efficient AI
luisdans · x · 2026-09-22
Microsoft's Luis Dans shared a video looking at the engineering behind efficient AI at scale: the team co-optimizes the entire stack around its in-house Maia 200 accelerator and Cobalt chips, from memory and precision formats to cooling and power delivery.
More from Infra
- GKE Pod snapshots cut AI inference cold starts by 89%, loading 70B models in 37s — rseroter · 2026-09-22
- Qwen 3.8 Flash Next hits 3.1k tok/s prefill on M5 Ultra — GabGarrett · 2026-09-22
- Wally inference stack debuts: GLM-5.3 Max at 790 tok/s for open models — ycombinator · 2026-09-22
- XGrammar-2 ships strict tool calling for agents, adopted by xAI, DeepSeek and vLLM — vllm_project · 2026-09-22
- vLLM offloads video decoding to NVDEC, 2x+ throughput on 8×H100 — vllm_project · 2026-09-22
- Bittensor subnet SOMA builds decentralized AI compression marketplace — markjeffrey · 2026-09-22