ByteDance Seed finds phase sensitivity in chunked KV-cache compression, retrieval accuracy swings 40 points
ByteDance-Seed · hf · 2026-10-01
A new study from ByteDance Seed exposes a hidden flaw in chunked KV-cache compression, a popular way to cut memory and attention costs for long-context inference.
- Compressing token windows at a fixed stride introduces a new positional coordinate: a token's phase, i.e. its position relative to compression-window boundaries. The same information can be easy to retrieve at one phase and hard at another.
- In large open-weight models using such compression, long-context retrieval accuracy can vary by up to 40 percentage points across phases — periodic weak spots that average benchmark scores conceal.
- The team pretrained transformers from scratch across multiple KV-compression designs and reproduced the effect in all variants. Causal-intervention mechanistic analysis shows attention components specialize asymmetrically by source phase, and idealized retrieval models suggest gradient flow dynamics may favor sharp phase specialization.
- Takeaway: evaluating models with chunked KV compression requires measuring across phases; high average accuracy can coexist with systematic positional failures.
More from Infra
- Cerebras COO sells $78M in stock to zero as OpenAI reportedly skips WSE chips — TheZachMueller · 2026-10-01
- Indie dev ternary-quantizes Qwen 27B to 7.3GB, keeps image skills for ComfyUI — udmrzn · 2026-10-01
- 8B model trained on distributed gaming GPUs for $6,500 runs on phone CPU at ~60 tok/s — markjeffrey · 2026-10-01
- New theory extends speculative decoding acceptance beyond distribution-preserving sampling — baseten · 2026-10-01
- ComfyUI launches Comfy API: deploy workflow JSONs as autoscaling endpoints — PurzBeats · 2026-10-01
- One prompt, 30 minutes: NVIDIA VSS Blueprint 3.3 builds production-line vision AI agents — NVIDIAAI · 2026-10-01