Tri Dao on the Next Big Unlock: Kernels, Inference Stacks, and Fleet Optimization

togethercompute · x · 2026-08-13

During an interview at ICML, FlashAttention creator Tri Dao shared his perspective on the evolution of AI architectures. When asked where the next major architectural unlock will come from, he offered a contrarian view: there isn't a single silver-bullet architecture.

Instead, he argues that the real performance unlocks lie in layer-by-layer optimizations—specifically in kernels, inference stacks, and fleet-level optimization. He also emphasized that most of this critical infrastructure work is currently happening in the open-source community.

Original post →

More from Infra

Infra channel →