DeepSeek's 890 bytes per token sparks demand for long-context evals
teortaxesTex · x · 2026-09-10
teortaxesTex highlights that DeepSeek's new model runs at an extremely low 890 bytes per token, and calls for long-context benchmark results to verify whether such aggressive KV-cache compression holds up on long documents.
Related event: DeepSeek compressed KV cache 54x in nine months, analysts say(4 posts)→
More from Models
- Same Bug Benchmark: GPT-6 Astra Medium Fixes 34/105, Low Scores 27 — PawelHuryn · 2026-09-10
- Bug Hunt Bench adds GPT-6 Astra scores: 27/105 low, 34/105 medium — PawelHuryn · 2026-09-10
- ChatGPT can't stop second-guessing you, and users blame its safety training — Due-Conference-5134 · 2026-09-10
- 105 Hidden Bugs Benchmark: DeepSeek V4.1 Flash Scores 24 at $1.80 vs Opus 5's $51 — sanderjson · 2026-09-10
- DeepSeek's answer to surging demand: make its model cheaper and faster — yacineMTB · 2026-09-10
- Huge share of post-2022 web data is AI content mislabeled as human-written — menhguin · 2026-09-10