Tri Dao on the Next Big Unlock: Kernels, Inference Stacks, and Fleet Optimization
togethercompute · x · 2026-08-13
During an interview at ICML, FlashAttention creator Tri Dao shared his perspective on the evolution of AI architectures. When asked where the next major architectural unlock will come from, he offered a contrarian view: there isn't a single silver-bullet architecture.
Instead, he argues that the real performance unlocks lie in layer-by-layer optimizations—specifically in kernels, inference stacks, and fleet-level optimization. He also emphasized that most of this critical infrastructure work is currently happening in the open-source community.
More from Infra
- Run a Local Real-time Voice AI Assistant in a Single Docker Container — tom_doerr · 2026-08-13
- Browserbase launches with funding to build a programmable browser for AI agents — jeff_weinstein · 2026-08-13
- Open-Sourced CUDA Programming Course Hits 3.9k Stars on GitHub — tom_doerr · 2026-08-13
- AMD Announces Day 0 Support for Qwen3.8-2.4T Open-Weight Model — AccBalanced · 2026-08-13
- Claude Opus 5 Tops InferenceBench with 8.90x Speedup via Adaptive Serving — maksym_andr · 2026-08-13
- Musk: SpaceX to Add 6-8GW Datacenters in 2027, Path to $300B ARR — AccBalanced · 2026-08-13