AI agent tunes Triton kernels on AMD MI210, flipping grid order yields 1.29x speedup
zmkzmkz · x · 2026-10-11
Developer Erland Hilman Fuadi shares how he used GEAK (an agent that writes and tunes GPU kernels) to optimize decode-phase GEMM on an AMD MI210, where M is tiny (1–64 tokens) but the weight matrix is large.
Results: weighted speedup went from 1.00x to 1.23x after round 1 and 1.29x after round 2. Round 1 produced expected optimizations: a Triton split-K kernel, retuned tile sizes, and a GEMV path for M=1.
Interesting finding: profiling showed M=4/16 already streamed weights near the card's ceiling (0.9–1.19 TB/s), but M=64 ran at only 625 GB/s because the tile was too large (BLOCKM=64, 150+ registers) and only 96 of 104 CUs were occupied.
Round-2 trick: switching to smaller M tiles, the agent flipped the order of programid axes in the Triton kernel so that M tiles of the same (N, K) weight block are adjacent in launch order, improving L2 reuse and cutting redundant weight reads.
More from coding & agent
- Dev hand-made 2 SwiftUI text animations, had Claude generate 182 more, open-sources all 184 — amos_gyamfi · 2026-10-11
- Ben Hylak on building simulations for agent evals: replay traces, detect sim awareness — HamelHusain · 2026-10-11
- Turingo detects AI writing by replaying document revision history, not text predictions — sethlazar · 2026-10-11
- gemini-cli VSCode extension leaked disposables due to comma-expression bug in activate() — nosmile99 · 2026-10-11
- Claude is an underrated used-car hunting tool: market modeling, scraping, daily alerts — eherrerosj · 2026-10-11
- OpenAI's Decisions API enters public beta, now callable via Apple's Foundation Model framework — rxwei · 2026-10-11