New HW Roofline Calculator Estimates Theoretical LLM Inference Speeds

Informal-Trouble2183 · reddit · 2026-09-28

A developer released a hardware roofline calculator that estimates theoretical decoding/prefill performance for LLMs based on model architecture, quantization, GPU and memory parameters. It's a theoretical bound, but useful as a step-0 sanity check for what fits on your hardware and how each bottleneck contributes. Available at ai-leaderboard.dev under "HW Roofline".

Original post →

More from Infra

Infra channel →