REFRAG beats prompt caching's exact-prefix limit, but 16x compression still won't fit your database in context

CShorten30 · x · 2026-10-05

In a podcast discussion on sparse attention and vector databases, CShorten30 explains: prompt caching reuses KV states only for an exact prefix, so reordering retrieved docs kills it; REFRAG instead precomputes one embedding per chunk that works in any prompt or order. But even 16x compression can't make "whole database in context" work — 1M docs at 500 tokens is 500M tokens, still 31M compressed. Search is here to stay.

Original post →

More from Infra

Infra channel →