Qwen 3.8 27b with DeepSeek Harness: 10h stable, 37-60 tok/s on RTX 3090

cviperr33 · reddit · 2026-08-16

A Reddit user shares experience with Qwen 3.8 27b paired with DeepSeek Harness (DSH). On an RTX 3090, the model ran for 10 hours without errors, auto-compressed context, stayed on task, and solved problems in one shot. Speeds: 37 tok/s with large context, 50 tok/s fresh, 56-60 tok/s in Unsloth Studio. Used UD q4 k xl quant + vision F16, 92k context, MTP + ngram, 66 GPU layers. User notes default xHigh thinking mode, sometimes thinking 20 minutes. Awaits 35b MoE model as dense model needs 16GB+ VRAM.

Original post →

More from Models

Models channel →