Modular shows unified compute layer: TPU v6e and d-Matrix Corsair brought up in days
carrycooldude · x · 2026-09-09
Modular chief scientist Abdul Dakkak's ModCon 2026 talk targets today's fragmented AI software stack — every accelerator ships its own compiler, kernels and runtime — with a unified compute layer and a repeatable recipe for bringing up new hardware:
- AWS Trainium: done by Modular's own team, running Gemma 4 31B end to end
- Google TPU v6e: brought up by partner HTEC with no LLVM backend to target, no prior Mojo/MAX knowledge, and minimal support
- d-Matrix Corsair: implemented by d-Matrix on their own stack — repo access Tuesday, working matmul by the following Monday
An accompanying HTEC post claims the approach cuts AI hardware enablement from years to months.
More from Infra
- Inference is turning GPU compute into a tradable commodity — ArtificialAnlys · 2026-09-09
- Cohere open-sources megakernel serving engine, up to 1.58x faster than vLLM — cohere · 2026-09-09
- Solving Navier-Stokes cost 130B output tokens — up to $18M depending on model pricing — mitsuhiko · 2026-09-09
- Baseten cuts delta weight syncs for frontier models to under 40 seconds — baseten · 2026-09-09
- Dell says DRAM, NAND shortages persist and nearly all leading-node products are constrained — Beth_Kindig · 2026-09-09
- Alphabet's CapitalG backs AI chip startup Celero at $3 billion valuation — dinabass · 2026-09-09