Kipply breaks down transformer inference arithmetic for H200/B200 in new perf engineering repo
ycombinator · x · 2026-09-29
waferai launched what it calls the most comprehensive AI performance engineering repo, with part 6 by Carol Chen (kipply) covering Transformer Inference Arithmetic.
- Builds an approximate inference cost model around per-token compute, bytes moved by the GPU, and inter-GPU communication, as a starting point for latency/throughput analysis on H200, B200, and B300.
- Examples use A100 GPUs, combined with NVIDIA Hopper and Blackwell documentation.
- Key takeaway: prefill and decode expose different token parallelism — prefill processes prompt tokens together for more weight reuse in projection/FFN matmuls, while small-batch decode can spend more time loading weights than computing. Profile each phase separately.
- Batching raises arithmetic intensity (thread truncated in source).
More from Infra
- Qwen3.8 Flash hits 74 tok/s single-stream, 212 tok/s aggregate on one DGX Spark — open vLLM recipe — DimeRhyme · 2026-09-29
- Redditor predicts sub-$1000 device running SOTA models will spawn the next big company — Robert__Sinclair · 2026-09-29
- Only 3 of ~6,000 data center projects hit by AI buildout moratoriums: SemiAnalysis — MatthewBerman · 2026-09-29
- Cloudflare birthday week ships 8 open source updates: forge, vinext 1.0, native Rust in Workers — ritakozlov · 2026-09-29
- Google Trends' #1 US region for every query is tiny Cheyenne, Wyoming — likely bot traffic — lilyraynyc · 2026-09-29
- Sentdex has run 4B+ tokens locally on GLM 5.3 Flash — his most-used local model ever — Sentdex · 2026-09-29