AI Performance Engineering resource list v2 covers everything from CUDA to MoE serving

AccBalanced · x · 2026-08-24

waferai released a major v2 update to its GPU and AI performance engineering resource list, billed as the most comprehensive resource for the field.

The opinionated list starts with how a single inference request works, then builds through the CUDA execution model, roofline analysis, transformer arithmetic, TTFT/TPOT/goodput, and kernel optimization. New additions include FlashAttention-4, Blackwell tensor memory and low-precision tensor cores; continuous batching, KV-cache systems, quantization, speculative and structured decoding; MoE serving, collectives, topology, and prefill/decode disaggregation; plus newer hardware (Blackwell Ultra, MI350/CDNA4, Ironwood, Trainium3) and benchmarks like Kernelbench-verified and SOL-execbench.

Original post →

More from Infra

Infra channel →