Building an ultra-high throughput AI-SQL engine: cutting LLM call costs with query plans
HamelHusain · x · 2026-10-08
A Full Stack Data Lab deep dive (Shreya Shankar, Charles Frye et al.) explains why AI-SQL — LLM-powered functions in Snowflake, BigQuery, Databricks, MotherDuck — gets prohibitively expensive when filters and joins trigger hundreds of thousands of LLM calls, and how query-plan-based optimization (building on DocETL, LOTUS, Palimpzest, ThalamusDB) cuts costs. Recommended by Hamel Husain as an approach for analyzing 6.5k pages of semi-sensitive documents, runnable on Modal.
Related event: Berkeley Team Details AI-SQL Engine That Slashes Millions of LLM Calls(2 posts)→
More from Infra
- Splash 1.3.0 cuts local agent first-token time from 19s to 1s via SSD offloading — songhan_mit · 2026-10-08
- StackOverflow 2026 Survey: Postgres Stays #1 at 58%, Supabase Climbs to 8% — dshukertjr · 2026-10-08
- Dev runs 456GB DeepSeek v4.1 on dual GPUs with 192GB VRAM, offloading experts to SSD — HankYeomans · 2026-10-08
- GB300 compute capacity bought with opencode credits in all-token deal — const_reborn · 2026-10-08
- Together is now the #1 provider by token volume on OpenRouter — NVIDIAAI · 2026-10-08
- umbrelOS 2.0 makes local AI a two-install setup: Ollama + Open WebUI with auto GPU detection — JosephJacks_ · 2026-10-08