Modular Claims Mojo Kernels Beat FlashAttention 4, Launches MAX Inference Framework
AI Engineer · youtube · 2026-10-12
At AI Engineer World's Fair 2026, Modular chief scientist Abdul Dakkak argued today's AI software stack is a mess — vLLM, SGLang, TensorRT-LLM and llama.cpp each patch around different hardware — while compute is booked out for years and new chips keep arriving.
Three pillars
- Mojo: a Pythonic systems language for CPUs, GPUs and accelerators aiming to match CUDA; open-sourced matmul and attention kernels he says beat vendor libraries and FlashAttention 4; Mojo 1.0 and open-source Mojo are coming
- MAX: an inference framework rethinking PyTorch-era assumptions, supporting thousands of model architectures behind OpenAI-compatible endpoints across GPUs, nodes and mixed hardware
- Modular Cloud ties it together as a "tunable glass box"
Benchmarks: Kimi K2.5, FLUX.2, AMD MI355 and Gemma 4 vs vLLM and torch.compile, plus sub-second image generation. Try it with pip install modular.
More from Infra
- AI buildout to cost $10.3 trillion to finance through 2032, topping all prior US investment booms — KyeGomezB · 2026-10-12
- Running a 456GB model on 192GB VRAM: offloaded inference hits 60-125 tok/s with 1M context — HankYeomans · 2026-10-12
- VitalOps launches agentic inference optimization, 2.6x median speedup — abhijithneil · 2026-10-12
- Running Qwen3.8 Flash-Next locally on AMD 7900 XTX at 500k context, 105-160 tok/s — human_in_the_looop · 2026-10-12
- Siemens Brings Nvidia Omniverse into Digital Twin Composer to Pave the Way for Physical AI — RevLebaredian · 2026-10-12
- OpenRouter token traffic explodes from 2T to 379T/month, open-weight models at 75% — Beth_Kindig · 2026-10-12