SALT: CELF-Based Sentence-Level Compression for KV Cache Retrieval
No_Sky9786 · reddit · 2026-08-19
The author introduces SALT, a structure using the CELF algorithm to retrieve only important sentences from the KV Cache. It shrinks long documents to a fixed size before sending them to the LLM, retaining the most informative sentences. Compatible with any model, it produces shorter prompts and cuts compute, memory, and latency. The project is open-source on GitHub. The author is seeking help to implement a dynamic budget adjustment method instead of a fixed 20-25% retrieval rate to optimize GPU prefill memory usage at scale.
More from Infra
- Developer Exhausts 20x Usage Limits on Both OpenAI and Anthropic — antirez · 2026-08-19
- Robotic Simulation Highlights as Clean Innovation Incentive System — const_reborn · 2026-08-19
- Users seek benchmarks on token subsidies for subscriptions like Claude and Codex Pro — mitsuhiko · 2026-08-19
- Meta Releases Muse Glimmer: 30B Model Optimized for Always-On Local Voice Agents — Once_ina_Lifetime · 2026-08-19
- Gortex Engine Indexes Code into Graphs to Cut Token Usage by 50x — tom_doerr · 2026-08-19
- A100 study shows high GPU utilization masks low tensor core efficiency — tokenbender · 2026-08-19