REVA Mines LLM Attention into Reusable Evidence Views, Cutting RAG Compression Overhead up to 15.6x
_reachsumit · x · 2026-09-11
REVA reframes RAG compression as data mining: instead of per-query compression, it aggregates historical query-document-model interactions into reusable evidence views. It mines the generator's past attention traces into a document-keyed, budget-agnostic score store, then renders budget-specific plain-text views that preserve document order.
Across four benchmarks and modern LLMs, REVA improves generation quality by 1.0-5.8 points over existing compressors while reducing compression overhead by 5.3-15.6x, adding under 40 ms of latency.
More from Infra
- Web Demo Approximates V4.1 Flash-Style Fast KV Prefill on Qwen3 — T_rex2700 · 2026-09-11
- OpenAI CFO: Compute I bought a year ago could sell for 3-5x today — and we're still short — rwang07 · 2026-09-11
- iFlytek's Spark X2.5 trained on 10,000 domestic Ascend 910B GPUs with 97% uptime — 机器之心 · 2026-09-11
- Edge0-35B-A3B preview MoE model for edge inference trends on Hugging Face — Edge0 · 2026-09-11
- Reflect Orbital wants to sell sunlight via volleyball-court mirrors on satellites — kyliebytes · 2026-09-11
- Data centers are for startups, not frontier labs: more compute is the anti-monopoly move — arthurcolle · 2026-09-11