Building an ultra-high throughput AI-SQL engine: cutting LLM call costs with query plans

HamelHusain · x · 2026-10-08

A Full Stack Data Lab deep dive (Shreya Shankar, Charles Frye et al.) explains why AI-SQL — LLM-powered functions in Snowflake, BigQuery, Databricks, MotherDuck — gets prohibitively expensive when filters and joins trigger hundreds of thousands of LLM calls, and how query-plan-based optimization (building on DocETL, LOTUS, Palimpzest, ThalamusDB) cuts costs. Recommended by Hamel Husain as an approach for analyzing 6.5k pages of semi-sensitive documents, runnable on Modal.

Related event: Berkeley Team Details AI-SQL Engine That Slashes Millions of LLM Calls(2 posts)→

Original post →

More from Infra

Infra channel →