Wafer launches comprehensive AI performance engineering repo covering CUDA, FlashAttention and quantization
Hesamation · x · 2026-09-21
Wafer has released what it calls the world's most comprehensive AI performance engineering repo, covering GPU fundamentals & CUDA, kernel optimization, FlashAttention, KV caching & quantization, plus NVIDIA, AMD and TPU architectures. All resources link to solid docs, papers and repos, and will be posted as an ongoing series — starting with NVIDIA's "Intro to CUDA C++" guide on kernel launches, thread indexing and memory management.
More from Infra
- Running 2.8T-param Kimi K3 on a 16x GB10 Cluster: 30 tok/s Coding Decode — ciprianveg · 2026-09-21
- Prompt cache restore works, but reusing KDA/Mamba states hits illegal memory access — TheZachMueller · 2026-09-21
- 128k context with compaction beats raw 256k/1M for quantized models — Informal-Trouble2183 · 2026-09-21
- US data centers to use more natural gas than Germany and Japan combined by 2035, says BNEF — Beth_Kindig · 2026-09-21
- RLinf + SGLang runs Cosmos3 at 3.33x end-to-end eval throughput for robots — ying11231 · 2026-09-21
- Use dirt-cheap Jev to pre-filter tasks before Claude and Codex to cut token costs — draginol · 2026-09-21