DeepSeek 发布 V4.1-Flash:552B MoE 仅 8B 激活,KV 缓存缩至 1/4

thione · x · 2026-09-21

DeepSeek 发布新架构家族中最小的模型 DeepSeek-V4.1-Flash:552B 参数 MoE,采用非对称 Causal Encoder–Decoder 架构,输入仅激活 8B 参数、输出 16B,带原生多模态视觉理解。

关键点:

所属事件:DeepSeek 发布 V4.1-Flash:552B MoE 架构为 Agent 重构(3 条相关)→

原文链接 →

「Infra」频道最新

更多「Infra」频道 AI 资讯 →