How to cost AI-powered filters: roofline model puts 5k-review LLM filter floor at 6.6s on H100
sh_reya · x · 2026-10-02
Shreya Shankar's team behind the open-source AI-SQL engine Quail explains how they estimate LLM query latency: instead of profiling per model/GPU/workload, they use a roofline model — divide arithmetic and memory traffic by GPU peak compute and bandwidth for a speed-of-light lower bound. The post covers transformer forward-pass costs, costing a single AI filter, ordering filter conjunctions, and an interactive Qwen3-4B-fp8 + H100 example. Verdict: the SoL for filtering 5k movie reviews is 6.6 seconds on one H100 — and no system today, including Quail, comes close.
More from Infra
- 180B Qwen model runs on one DGX Spark: 2.39-bit quant keeps 95.5% of BF16 scores — TheZachMueller · 2026-10-02
- Google: Starship must launch 1,600 times before space data centers work — TechCrunch AI · 2026-10-02
- Open-Source Local AI Avatar App Uses LM Studio and Chatterbox Voice Cloning — TheRedHairedHero · 2026-10-02
- Big Tech's AI Capex Now Accounts for Half of Wall Street's Profit Growth — speckx · 2026-10-02
- Hugging Face cofounder launches a million sandboxes live on stage at Modal Runtime — graceisford · 2026-10-02
- Uno Speculative Decoding Hits 1.30x on 28k-Token Prompts, 2x DFlash, Now in vLLM — yuntiandeng · 2026-10-02