DeepSeek's 890 bytes per token sparks demand for long-context evals

teortaxesTex · x · 2026-09-10

teortaxesTex highlights that DeepSeek's new model runs at an extremely low 890 bytes per token, and calls for long-context benchmark results to verify whether such aggressive KV-cache compression holds up on long documents.

Related event: DeepSeek compressed KV cache 54x in nine months, analysts say(4 posts)→

Original post →

More from Models

Models channel →