How to cost AI-powered filters: roofline model puts 5k-review LLM filter floor at 6.6s on H100

sh_reya · x · 2026-10-02

Shreya Shankar's team behind the open-source AI-SQL engine Quail explains how they estimate LLM query latency: instead of profiling per model/GPU/workload, they use a roofline model — divide arithmetic and memory traffic by GPU peak compute and bandwidth for a speed-of-light lower bound. The post covers transformer forward-pass costs, costing a single AI filter, ordering filter conjunctions, and an interactive Qwen3-4B-fp8 + H100 example. Verdict: the SoL for filtering 5k movie reviews is 6.6 seconds on one H100 — and no system today, including Quail, comes close.

Original post →

More from Infra

Infra channel →