Fleet: Multi-Die GPU Task Abstraction

matt_d · hn · 2026-07-16

This paper proposes Fleet, a megakernel hierarchical task abstraction designed for multi-die GPUs. The title itself indicates a focus on system-level optimization for GPU execution and task organization, rather than model or application-layer content.

In terms of direction, it aligns more closely with AI infrastructure: investigating how to more effectively organize and run large-scale kernels on new GPU form factors, providing foundational methods for future training/inference system optimizations.

Original post →

More from Infra

Infra channel →