Qwen3.8-27B hits ~40 tok/s on dual RTX 3060, passes CUDA coding test

anderspitman · reddit · 2026-08-15

Reddit user anderspitman shares early performance report of Qwen3.8-27B on 2x RTX 3060 12GB. With specific llama.cpp config, 40 tok/s achieved. In a CUDA vector add smoke test, the model one-shot generated the program and installed compiler, but couldn't run due to VRAM limits. It attempted debugging, and eventually succeeded after shutting down llama-server.

Original post →

More from coding & agent

coding & agent channel →