Nebius/WEKA 实测:分布式 KV cache 让单节点 agent 推理吞吐提升 2.4 倍

AccBalanced · x · 2026-09-25

Nebius 与 WEKA 在 NVIDIA HGX B300 上对共享 KV cache 层(WEKA NeuralMesh Augmented Memory Grid)做了 8 小时持续压测,重放真实 agentic coding 流量运行 DeepSeek-V4-Pro。

原文链接 →

「Infra」频道最新

更多「Infra」频道 AI 资讯 →