Meta unveils MTIA 300, its first training chip with built-in NIC chiplets
Meta_Engineers · x · 2026-08-31
Meta's engineering blog details MTIA 300, the first of its in-house accelerator family optimized for training ranking and recommendation models.
- Unlike LLMs, recommendation models are communication-bound: embedding tables can hold over 99% of parameters, driving frequent AllReduce/AllToAll/AllGather collectives across hundreds of accelerators.
- On GPUs, these collectives compete with compute for resources, leaving hardware underutilized.
- MTIA 300's built-in NIC chiplets and communication-offloading engines make communication a first-class citizen, beating general-purpose GPUs on these workloads.
- Meta co-designed the HCCL communication library alongside the chip.
Related event: Meta Unveils MTIA 300, Its First Training-Focused AI Accelerator(2 posts)→
More from Infra
- TensorSharp vs llama.cpp: Qwen 3.8 Flash Next Benchmarks — fuzhongkai · 2026-09-01
- Why did increasing context size increase speed in Llama.cpp? — satnl · 2026-09-01
- AI inference demand surges again, supply brutally outpaced by token growth — Baconbrix · 2026-09-01
- Warp founder predicts cloud-based collaborative factories for all companies within a year — charlieholtz · 2026-09-01
- JPM: 1GW of AI Infrastructure Costs $40-45B, Frontier Labs Make ~$30B per GW — zephyr_z9 · 2026-09-01
- Data Center Worker: Fastest Blue-Collar Path to Six Figures Right Now — AICopyLab · 2026-09-01