ByteDance Seed Paper Explains Phase Blind Spots in KV Compression Behind DeepSeek's Erratic Long- Context Performance
teortaxesTex · x · 2026-10-08
A new paper from ByteDance's Seed team explains why DeepSeek's long-context performance seems to come and go: chunked KV compression introduces "phase blind spots," and where key information falls within a compression window determines how easy it is to retrieve.
- The fluctuation period exactly matches the compression step: V4 uses a CSA step of 4, so capability wobbles every 4 tokens; V4.1 changed the step to 2, and the period became 2 as well
- Extreme case: when asked to complete FP8 inference code, V4-Flash-Base flips between correct FP8 and wrong FP32 with a 4-token period just from changing the length of a leading comment
- In 128K needle-in-a-haystack tests, the same "needle" scored up to 40 percentage points differently depending on position, with average scores masking weak phases
- The phenomenon appears in the post-trained V4-Flash, the larger V4-Pro, and V4.1 alike — a structural side effect of chunked compression, not an occasional bug
- Takeaway: without sweeping across phases, long-context benchmarks systematically overstate scores
More from Models
- SentenceTransformers gets native ColPali model support thanks to tomaarsen — tomaarsen · 2026-10-08
- ChatGPT Plus users report 'thinking' time doubled in 2025 with no quality gain — Gazialp · 2026-10-08
- Mistral Large 4 debuts at #45 on Code Arena WebDev, near Opus 4.8 at 1/6 the price — arena · 2026-10-08
- Liquid AI releases Open d1: open-weight 3B and 600M multimodal decision models — JosephJacks_ · 2026-10-08
- Liquid AI details d1-omni-600M: 600M params for text+image or text+audio — JosephJacks_ · 2026-10-08
- Anthropic Staffer: Opus 3 Doesn't Follow Our Constitution, But It Saw the Sincerity — repligate · 2026-10-08