DEX-Comp Compresses RAG Context 16x While Matching or Beating Uncompressed Baselines
_reachsumit · x · 2026-09-07
A new arXiv paper proposes DEX-Comp, a two-stage training recipe for soft context compression in RAG that breaks past the ceiling of distillation-only approaches:
- Pure Distillation: warm-starts the compression model via imitation on correct responses from uncompressed RAG only;
- Hard Exploration: runs reinforcement learning solely on queries where uncompressed RAG fails, pushing the model toward computation patterns better suited to compressed representations.
Across five open-domain QA benchmarks at retrieval depths from top-5 to top-30, DEX-Comp compresses retrieved contexts 16x, speeds up inference 4x–24x, and matches or exceeds the uncompressed RAG baseline. Ablations confirm each stage's contribution and generalization across datasets and backbones.
More from Infra
- Redditor builds local agent 'Sai' with a real body, thermal senses and skin in the game — Quebber · 2026-09-07
- One interconnect session, four competing standards: 'optical interconnects are a filthy mess' — jwt0625 · 2026-09-07
- PyTorch Foundation in Shanghai: 80M monthly downloads, tackling diverse-hardware open stack — PyTorch · 2026-09-07
- Embedding Surgery: query-time localized vector edits fix dense retrieval rankings, +60% nDCG@10 — _reachsumit · 2026-09-07
- VDN-H3 ported to Vpipe hits ~2.6x faster H3 generation on M5 Pro 24GB — TgoAI · 2026-09-07
- SC Asia 2027 opens call for papers on supercomputing and AI infra, due Oct 7, 2026 — thoefler · 2026-09-07