Wafer launches comprehensive AI performance engineering repo, starting with deep dive on Transformer inference
garrytan · x · 2026-09-12
Wafer.ai launched what it bills as the most comprehensive AI performance engineering resource repo, retweeted by Garry Tan, who says fully understanding the first resource puts you ahead of 90% of people on inference.
Part 1, "All About Transformer Inference" (from How To Scale Your Model), covers:
- arithmetic intensity of linear layers and attention across prefill and decode
- KV cache sizing by layer count, KV heads, head dimension, sequence length, and precision
- the compute/HBM bandwidth crossover and how batch size and quantization shift it
- decode latency and serving decisions
More resources from the repo will be posted sequentially.
Related event: Wafer Launches Comprehensive AI Performance Engineering Resource Library(2 posts)→
More from coding & agent
- Agents need explicit explain/propose/execute modes: 'how would you' is a proposal, not permission — OriginalHospital · 2026-09-12
- Scaling parallel Claude Code agents: the bottleneck is now human oversight, not the model — tylerjdunn · 2026-09-12
- Dev burns entire Codex quota in under 56 minutes after reset — chriscarm · 2026-09-12
- Pipecat v1.9 adds Meta's Muse Voice Transcribe, the lowest semantic-WER STT model tested — solyarisoftware · 2026-09-12
- Jensen Huang: Even AGI Won't Auto-Boost Company Productivity — You Need an Agent Harness — 大模型之路 · 2026-09-12
- How to Build a 'Company Brain' That Gets Smarter Every Week — Roger_M_Taylor · 2026-09-12