New HW Roofline Calculator Estimates Theoretical LLM Inference Speeds
Informal-Trouble2183 · reddit · 2026-09-28
A developer released a hardware roofline calculator that estimates theoretical decoding/prefill performance for LLMs based on model architecture, quantization, GPU and memory parameters. It's a theoretical bound, but useful as a step-0 sanity check for what fits on your hardware and how each bottleneck contributes. Available at ai-leaderboard.dev under "HW Roofline".
More from Infra
- Fireworks' Ember-1 post-trains Kimi K3 to reason 40% more concisely at same quality — isidentical · 2026-09-28
- One AMD driver flag boosts dual-GPU Vulkan LLM inference up to 4x — tabletuser_blogspot · 2026-09-28
- PrismML's Bonsai 2 shrinks Qwen3.8 27B to 5.9GB, keeps 98.2% capability, runs on a 5090 — dl_weekly · 2026-09-28
- Developer Plans to Let Codex Pick Which Tests Run, Slashing CI Costs — sull · 2026-09-28
- Free online guide covers LLMs from first principles to local deployment — JFPuget · 2026-09-28
- Is a vector database enough for production AI agents? Reddit debates storage design — OkShirt9372 · 2026-09-28