Qwen3.8-27B在双RTX 3060上跑出40 tok/s,实测CUDA编程能力

anderspitman · reddit · 2026-08-15

Reddit user anderspitman shares early performance report of Qwen3.8-27B on 2x RTX 3060 12GB. With specific llama.cpp config, 40 tok/s achieved. In a CUDA vector add smoke test, the model one-shot generated the program and installed compiler, but couldn't run due to VRAM limits. It attempted debugging, and eventually succeeded after shutting down llama-server.

原文链接 →

「编程与Agent」频道最新

更多「编程与Agent」频道 AI 资讯 →