Batch AI filtering done right: Qwen3-4B on one H100 filters 5k movie reviews in ~6.6 seconds
HamelHusain · x · 2026-10-02
Hamel Husain amplified shreya's point that calling an API once per row for batch decision tasks is a terrible idea — you get no benefit from query planning and sit extremely far from optimal performance.
- Core argument: run small local models on GPU for batch filtering instead of per-row cloud API calls
- Hard numbers: with Qwen3-4B on a single H100, the speed-of-light for filtering 5,000 movie reviews is about 6.6 seconds
- The 'batch decision model + query planning' approach is highly relevant for large-scale labeling and content filtering pipelines
Related event: Quail Team Shows Batch LLM Inference Beats Row-by-Row API Calls(2 posts)→
More from coding & agent
- Making a one-video history of the internet with Claude inside Cursor — prasenx · 2026-10-02
- Opus 5.5 turns Strudel live coding into an interactive jam session via MCP — repligate · 2026-10-02
- Skyvern 3.0 rebuilt from scratch hits SOTA 90.5% on Odyssey browser agent benchmark, stays open source — ycombinator · 2026-10-02
- How to test model switches for AI agents before shipping to catch silent tool-calling regressions — Fun_Employment6042 · 2026-10-02
- Databricks launches ai_decide, a sub-second structured decision AI function cheaper than LLMs — matei_zaharia · 2026-10-02
- Ivo's contract agent hits 91% on Legal Agent Benchmark via River AI post-training — ibab · 2026-10-02