Modular Claims Mojo Kernels Beat FlashAttention 4, Launches MAX Inference Framework

AI Engineer · youtube · 2026-10-12

At AI Engineer World's Fair 2026, Modular chief scientist Abdul Dakkak argued today's AI software stack is a mess — vLLM, SGLang, TensorRT-LLM and llama.cpp each patch around different hardware — while compute is booked out for years and new chips keep arriving.

Three pillars

Benchmarks: Kimi K2.5, FLUX.2, AMD MI355 and Gemma 4 vs vLLM and torch.compile, plus sub-second image generation. Try it with pip install modular.

Original post →

More from Infra

Infra channel →