DeepSeek 发布 V4.1-Flash KV 缓存压缩法,大幅降低推理内存成本

justlikemedics · reddit · 2026-09-13

Reddit 用户分析称 DeepSeek 随 DeepSeek-V4.1-Flash 发布了一种大幅压缩 KV cache 内存占用的新方法。

所属事件:DeepSeek V4.1-Flash 将 KV 缓存压至每 token 890 字节(2 条相关)→

原文链接 →

「创投」频道最新

更多「创投」频道 AI 资讯 →