Dual 3090 Qwen 27B Full Context Setup: 3200 tok/min Achieved
CryptographerLow7817 · reddit · 2026-08-23
Author shared a production-ready setup for running Qwen3.8-27B with 262k context on dual RTX 3090s (no NVLink). Using vLLM 0.27.1, MTP-3 speculative decoding, and GPU prefix caching, a 96.8% cache hit rate was achieved. Performance: Mean TTFT 6.5s (vs 140s cold), single-session decode at 80-90 tok/s, and multi-session throughput of 3200 tok/min.
More from coding & agent
- Agent hallucinations spread like wildfire; dev fixes it by disabling history saves — Vjeux · 2026-08-23
- Agent swarms prone to hallucination cascades and infinite loops — Vjeux · 2026-08-23
- Dev shares 'reality-check' agent workflow for project review — doodlestein · 2026-08-23
- Microsoft releases LangChain.js for Beginners course — adnan_hashmi · 2026-08-23
- Seeking Best Value AI Agent + Model Setup as DeepSeek Prices Rise — ProudCordonian · 2026-08-23
- Dev bypasses Claude copyright refusal with DeepSeek V4 Pro for movie plugin — AlchainHust · 2026-08-23