Qwen3.8-27B在双RTX 3060上跑出40 tok/s,实测CUDA编程能力
anderspitman · reddit · 2026-08-15
Reddit user anderspitman shares early performance report of Qwen3.8-27B on 2x RTX 3060 12GB. With specific llama.cpp config, 40 tok/s achieved. In a CUDA vector add smoke test, the model one-shot generated the program and installed compiler, but couldn't run due to VRAM limits. It attempted debugging, and eventually succeeded after shutting down llama-server.
「编程与Agent」频道最新
- Teknium 称自己被 Agent 杠杆率过高 — nickbaumann_ · 2026-08-15
- GraphJin宣称用小型廉价模型驱动整个组织,Agent harness成关键 — dosco · 2026-08-15
- Anthropic 分享构建高性价比 Agent 的技巧 — brada · 2026-08-15
- Actual 即将集成 Hermes:开箱即用的本地 Agent 工作站 — markjeffrey · 2026-08-15
- Agora:专为AI智能体打造的文本开放世界游戏,通过MCP协议交互 — Own_Assistant_2511 · 2026-08-15
- 修复并改进 Qwen 3.8 聊天模板:解决推理与工具调用问题 — Chromix_ · 2026-08-15