Open-source AI-SQL engine Quail adds prefix sharing: 3x fewer tokens, 2.6x faster queries
sh_reya · x · 2026-10-08
shreya announced that Quail, an open-source AI-SQL engine, now supports prefix sharing. In a real example tagging 1,772 long agent traces with qwen3-4b, it computes 3x fewer tokens and runs 2.6x faster.
Thread highlights:
- Same idea as prefix/radix caching in inference engines, but since LLM calls arrive via AI-SQL queries, sharing is planned upfront during query optimization — independent of request arrival order or KV cache eviction
- Three reuse forms: KV rewind, prefix sharing, and KV retention across operators
- KV stored in 16-token pages; multiple documents can reference the same page
- Enabled as a cost-based optimizer rule: Quail weighs document overlap against KV write cost automatically
- Cost model estimates GPU time from compute and memory bandwidth, verifiable via explain(analyze=True)
Related event: Open-Source AI-SQL Engine Quail Adds Prefix Sharing(2 posts)→
More from Infra
- Windows demo routes coding tasks to local model with GPU spinning, llama.cpp lands on Windows ML — ryanshrout · 2026-10-08
- exe.dev explains crossing the hyper-thread boundary: core scheduling cookies for VM isolation — davidcrawshaw · 2026-10-08
- NVIDIA's LoGRA cuts RL training memory by up to 45.7%, trains 27B model where Adam OOMs — mark_k · 2026-10-08
- Qwen3.8-Flash-Next on 6x3090 without NVLink: prefill 8-10x faster, long-context decode 2-3x — flynth92 · 2026-10-08
- Google DeepMind's Philipp Schmid: Give Every AI Agent Its Own Managed Cloud Sandbox — AI Engineer · 2026-10-08
- GitHub goes down; engineers confirm a fix is in progress — iannuttall · 2026-10-08