Running Qwen3.8-27B on dual RTX 3060s: 50 tok/s recipe
Ecstatic-Wash-7667 · reddit · 2026-08-15
A Reddit user shares a llama.cpp configuration for running Qwen3.8-27B on dual RTX 3060 12GB, achieving 50.6 tok/s generation with MTP speculative decoding, 573 tok/s prefill, and 23.2/24.0 GiB VRAM usage.
Related event: Qwen3.8-27B Hits 50 tok/s on Dual RTX 3060s(2 posts)→
More from coding & agent
- Cloudflare open-sources its AI productivity environment, Cloudflare OS — hichaelmart · 2026-08-15
- Real-world example: Multi-agent decision making for test optimization — iamrobotbear · 2026-08-15
- Designing Payment Authorization for AI Agents: Balancing Security and Autonomy — NoCalendar831 · 2026-08-15
- Ex-Meta Scientist: Agents should use the web like humans via pixels and clicks — DhruvBatra_ · 2026-08-15
- OpenAI Codex Error: 'gpt-5.6-sol' Model Possibly Deprecated — AKsnipebuster47 · 2026-08-15
- reBot Arm Control Stack Integrates Agentic AI, VLM, and LLM for Robotics — kamathsblog · 2026-08-15