Building LLM Inference CUDA Kernels From Scratch

Abhishekcur · x · 2026-07-13

The author shares their experience building LLM inference kernels from scratch using CUDA to truly understand the compute characteristics of the decode phase.

What was done

What was measured

The author also mentions that their favorite part is the content related to flash-decoding.

Related event: Developer Builds LLM Inference CUDA Kernels from Scratch(3 posts)→

Original post →

More from coding & agent

coding & agent channel →