30 Days of Inference: Deep dive into Blackwell and CUDA
blelbach · x · 2026-08-26
Developer Anirudh is launching a "30 Days of Inference" challenge, pushing a Blackwell DGX Spark to its limits. The curriculum covers Transformer architecture under the hood, Attention mechanisms and custom kernel implementations, CUDA Kernels (C++ & CuTile DSL), Kernel Profiling (Nsight), GPU Hardware specifics (Warps, SMs, Blackwell architecture), Inference Bottlenecks (Batching, Roofline), and custom Transformer development. Notes and resources will be shared daily, based on Philip Kiely's book and NVIDIA documentation.
More from Infra
- Alibaba's RecGPT-Mobile-V2: On-Device Behavior Prediction with RL — _reachsumit · 2026-08-26
- AMD MI350X Runs Qwen3.6-35B: Open Source Kernel Achieves 78.5k tok/s on 8 GPUs — SmilingGen · 2026-08-26
- MetricFire releases MCP server to query monitoring data with AI tools — PKMNPinBoard · 2026-08-26
- How continuous batching keeps GPUs busy: LLM inference runs steps, not requests — arpit_bhayani · 2026-08-26
- OpenAI's new chip allegedly 2x better perf/watt than Nvidia's Rubin — nickbaumann_ · 2026-08-26
- Frequent Claude outages in August spark user concerns over model degradation — 新智元 · 2026-08-26