Stanford's Prefix Sliding cuts long-reasoning inference ~3x without retraining, AIME25 score intact
rohanpaul_ai · x · 2026-09-04
A new Stanford paper shows long chain-of-thought doesn't need full in-memory attention: Prefix Sliding keeps the fixed task/tool prefix plus a sliding window of recent tokens, dropping intermediate reasoning.
- Attention analysis shows the prefix and latest tokens get nearly all attention weight.
- On Qwen3-1.7B, a 4,096-token window scores 33.9% on AIME25 vs 34.2% full attention, 3x faster inference, no retraining.
- Constant per-token cost once the window fills also enables RL rollouts beyond 100,000 tokens.
- Beats pure sliding windows, repeated summarization, and last-k deletion on speed/accuracy tradeoffs.
Paper: "Prefix Sliding for efficient test-time scaling"
More from Infra
- Why Gated DeltaNet survives 4-bit: full NVFP4 W4A4 quantization of a hybrid 27B LLM — minima-ai · 2026-09-04
- Anthropic's IREN deal: pricing and compute type matter more than the customer name — Kyrannio · 2026-09-04
- ByteDance to Add 5-6GW of AI Compute in Inner Mongolia, Costing Up to $135B — pstAsiatech · 2026-09-04
- ASML relies heavily on FPGAs, and Xilinx once caused major backlogs — PAstynome · 2026-09-04
- Tenstorrent Launches JapanFold: Sovereign Open-Source Drug Discovery for Japan — DavidBennett__ · 2026-09-04
- Analyst: no model trained without NVIDIA has ever beaten one trained on NVIDIA — BenBajarin · 2026-09-04