Amazon Paper: Training Should Match KV-Cache Inference Strategies

Amazon's new paper argues that KV-cache strategies should shape training, not just inference. Since sparse attention causes models to forget context during long-context inference, training should simulate this forgetting, which the paper shows mitigates long-context degradation.

2026-08-31 ~ 2026-08-31 · 2 related posts