Stanford's Prefix Sliding cuts long-reasoning inference ~3x without retraining, AIME25 score intact

rohanpaul_ai · x · 2026-09-04

A new Stanford paper shows long chain-of-thought doesn't need full in-memory attention: Prefix Sliding keeps the fixed task/tool prefix plus a sliding window of recent tokens, dropping intermediate reasoning.

Paper: "Prefix Sliding for efficient test-time scaling"

Original post →

More from Infra

Infra channel →