Prefix Sliding technique boosts LLM inference speed by 3x
iScienceLuvr · x · 2026-08-27
Research shows that most intermediate reasoning tokens lose importance over time, questioning the cost of retaining them. The proposed 'Prefix Sliding' method discards irrelevant tokens during inference. It achieves 3x speedup without training and better performance with RL training for traces over 100k tokens.
Related event: Prefix Sliding Speeds Up Inference 3x by Dropping Stale Tokens(2 posts)→
More from Infra
- A $6,000 rig to run full, non-distilled DeepSeek-R1 locally — carrigmat · 2026-08-27
- Recompile, Test, Run: Frontier Intelligence Locally, Ready for 10T Models by 2027 — carrigmat · 2026-08-27
- Speculatively Prefetch MoE Experts: Start from llama.cpp PR #25294 — carrigmat · 2026-08-27
- llama.cpp Underuses Your NVMe Array? The Author Lets Kimi Tweak the Code — carrigmat · 2026-08-27
- 16 Cheap 1TB Gen5 Drives Cost About $3,500 — Enough for Local Frontier Models — carrigmat · 2026-08-27
- No 16 NVMe Ports? Every PCIe Slot Splits Into Four of Them — carrigmat · 2026-08-27