REFRAG beats prompt caching's exact-prefix limit, but 16x compression still won't fit your database in context
CShorten30 · x · 2026-10-05
In a podcast discussion on sparse attention and vector databases, CShorten30 explains: prompt caching reuses KV states only for an exact prefix, so reordering retrieved docs kills it; REFRAG instead precomputes one embedding per chunk that works in any prompt or order. But even 16x compression can't make "whole database in context" work — 1M docs at 500 tokens is 500M tokens, still 31M compressed. Search is here to stay.
More from Infra
- UBS raises Nvidia 2027 GPU forecast by 600K units to 8.8M on Rubin ramp — Beth_Kindig · 2026-10-05
- Tesla's per-car inference compute scaling was unprecedented — and today's compute shortage is just the tip — yunta_tsai · 2026-10-05
- a16z: Old GPUs are turning into appreciating assets, per real rental and residual value data — ns123abc · 2026-10-05
- Fully local parkour sim vibe-coded with GLM 5.3 Flash on 2x DGX Sparks, recipe open-sourced — -dysangel- · 2026-10-05
- Budget Vulkan GPU list for local AI: used MI50 16GB at ~$150 is the top pick — tabletuser_blogspot · 2026-10-05
- Uber details its MCP Gateway: 800+ MCP servers, 5,000+ tools, auto-generated via AutoCrawler — Roger_M_Taylor · 2026-10-05