单张 RTX 5090 跑 Qwen3.8 Flash Next:靠 expert caching 达 50 t/s

dir3ctly · reddit · 2026-09-20

一位 Reddit 用户分享用 FreeToken 在单张 RTX 5090 上本地运行 Qwen3.8 Flash Next 的实测:文本生成约 50 t/s(40-60 浮动),prefill 约 2300 t/s,长上下文下依然稳定。

原文链接 →

「Infra」频道最新

更多「Infra」频道 AI 资讯 →