AI Performance Engineering resource list v2 covers everything from CUDA to MoE serving
AccBalanced · x · 2026-08-24
waferai released a major v2 update to its GPU and AI performance engineering resource list, billed as the most comprehensive resource for the field.
The opinionated list starts with how a single inference request works, then builds through the CUDA execution model, roofline analysis, transformer arithmetic, TTFT/TPOT/goodput, and kernel optimization. New additions include FlashAttention-4, Blackwell tensor memory and low-precision tensor cores; continuous batching, KV-cache systems, quantization, speculative and structured decoding; MoE serving, collectives, topology, and prefill/decode disaggregation; plus newer hardware (Blackwell Ultra, MI350/CDNA4, Ironwood, Trainium3) and benchmarks like Kernelbench-verified and SOL-execbench.
More from Infra
- xllm generates an image in 0.4 seconds — warycat · 2026-08-24
- WULF CEO reveals modern AI data centers use minimal water via closed-loop systems — robleclerc · 2026-08-24
- Cursor Team Publishes 'Git at Any Scale', Advocating for Stateless Infrastructure — thesephist · 2026-08-24
- Semiconductor engineers now more prestigious than doctors in South Korea — SuB8u · 2026-08-24
- Local development is fast and controllable, why rely solely on the cloud? — vboykis · 2026-08-24
- Trained on PrimeIntellect, rollouts rendered on Modal — willcb · 2026-08-24