Seeing like a GPU: axis diagrams make attention's parallelization patterns obvious

vtabbott_ · x · 2026-09-28

A follow-up from vtabbott and Gioele Zardini's updated diagrams work: GPUs see everything in terms of axes — batching, parallelizing, streaming, and reducing all happen over axes.

Redrawing attention with axis diagrams makes it immediately visible how the q-axis is parallelized in each expression, information the usual attention notation obscures — such as which axis the matmul applies over and which axis softmax reduces. The approach lets you 'see like a GPU' and spot parallelization patterns at a glance.

Related event: vtabbott_ and MIT's Gioele Zardini Ship Major Update to Interactive Attention Diagrams(5 posts)→

Original post →

More from Infra

Infra channel →