DeepSeek's V4.1 paper reveals KV cache reuse across checkpoints — 'a very strange model factory'
teortaxesTex · x · 2026-09-10
Digging into DeepSeek's V4.1 paper, Astra Pro surfaced several intriguing details — most notably that DeepSeek reuses KV caches between checkpoints. Commentator teortaxesTex calls the practice highly unusual, saying DeepSeek has 'a very strange model factory going,' hinting at an unconventional training pipeline.
More from Research
- Robotic fly achieves straight, level flight and recovers from wind gusts — moyix · 2026-09-11
- Two Millennium Problems solved within a month? Viral visualization gets a clearer remake — lishali88 · 2026-09-10
- AWS introduces AEM, a turn-level metric to isolate cascading agent errors — AWS ML Blog · 2026-09-10
- NVIDIA open-sources BioNeMo Inference Runtime, the GPU engine behind AlphaFold DB at million-scale — AllThingsApx · 2026-09-10
- Perceptual Losses paper wins ECCV Test of Time Award, still core to diffusion VAEs — CSProfKGD · 2026-09-10
- Amazon Science: why machine learning research agents don't overfit—and what compression has to do with it — Aaroth · 2026-09-10