Qwen3.8-27B hits ~40 tok/s on dual RTX 3060, passes CUDA coding test
anderspitman · reddit · 2026-08-15
Reddit user anderspitman shares early performance report of Qwen3.8-27B on 2x RTX 3060 12GB. With specific llama.cpp config, 40 tok/s achieved. In a CUDA vector add smoke test, the model one-shot generated the program and installed compiler, but couldn't run due to VRAM limits. It attempted debugging, and eventually succeeded after shutting down llama-server.
More from coding & agent
- Free Claude Code Highlights Value of Local Routing Layer — JeremyCMorgan · 2026-08-15
- MCP Stateless Protocol Revision Admits Original Flaws — DavidLinthicum · 2026-08-15
- AgenticROS Skill JARVIS Enables Voice and Vision Interaction — chrismatthieu · 2026-08-15
- Open source novel-script completes the industrialized AI short drama pipeline — EAccelerate_42 · 2026-08-15
- Qwen Agent Creates Process to Screenshoot and Test Its Own Game — OkFly3388 · 2026-08-15
- Teknium says he is over-leveraged by agents — nickbaumann_ · 2026-08-15