AI agent tunes Triton kernels on AMD MI210, flipping grid order yields 1.29x speedup

zmkzmkz · x · 2026-10-11

Developer Erland Hilman Fuadi shares how he used GEAK (an agent that writes and tunes GPU kernels) to optimize decode-phase GEMM on an AMD MI210, where M is tiny (1–64 tokens) but the weight matrix is large.

Results: weighted speedup went from 1.00x to 1.23x after round 1 and 1.29x after round 2. Round 1 produced expected optimizations: a Triton split-K kernel, retuned tile sizes, and a GEMV path for M=1.

Interesting finding: profiling showed M=4/16 already streamed weights near the card's ceiling (0.9–1.19 TB/s), but M=64 ran at only 625 GB/s because the tile was too large (BLOCKM=64, 150+ registers) and only 96 of 104 CUs were occupied.

Round-2 trick: switching to smaller M tiles, the agent flipped the order of programid axes in the Triton kernel so that M tiles of the same (N, K) weight block are adjacent in launch order, improving L2 reuse and cutting redundant weight reads.

Original post →

More from coding & agent

coding & agent channel →