DeepSeek's V4.1 paper reveals KV cache reuse across checkpoints — 'a very strange model factory'

teortaxesTex · x · 2026-09-10

Digging into DeepSeek's V4.1 paper, Astra Pro surfaced several intriguing details — most notably that DeepSeek reuses KV caches between checkpoints. Commentator teortaxesTex calls the practice highly unusual, saying DeepSeek has 'a very strange model factory going,' hinting at an unconventional training pipeline.

Related event: DeepSeek V4.1 Reportedly Returns to Encoder-Decoder Architecture, Inspired by YoCo(5 posts)→

Original post →

More from Research

Research channel →