Qwen3-Coder 30B本地跑进GitHub Copilot,96GB显存下每秒85 tokens

ollama · x · 2026-08-13

开发者burkeholland在GitHub Copilot应用中通过Ollama本地运行Qwen3-Coder 30B模型,在96GB VRAM上达到约85 tokens/秒的生成速度。虽然不及Sol,但进展可观。

原文链接 →

「编程与Agent」频道最新

更多「编程与Agent」频道 AI 资讯 →