Nemotron 3.5 Lightning 实测:并发吞吐提升16倍延迟不降

rhythmrg · x · 2026-08-11

Applied Compute 平台宣布支持 NVIDIA 的 Nemotron 3.5 Lightning 模型用于训练和推理。在智能体编码基准测试中,当并发量和总 token 吞吐量扩展 16 倍时,解码吞吐量、首 token 响应时间以及中位用户延迟基本保持不变。这得益于其 LatentMoE 和 Mamba 架构,能够以极低的开销扩展稀疏性、上下文长度和批处理大小。

所属事件:Nemotron 3.5 Lightning实测主打极速推理(2 条相关)→

原文链接 →

「Infra」频道最新

更多「Infra」频道 AI 资讯 →