Dual 3090 Qwen3.8-27B: DFlash2 Speed Test and Full Config Guide
Old_Ad_6033 · reddit · 2026-08-29
A Redditor shared a full recipe for running the Qwen3.8-27B INT4 model on dual RTX 3090s (24GB). Using a stack of vLLM 0.28.0, LMCache, and DFlash2, the setup leverages NVMe as a Level 2 KV cache to support 256k context length. Benchmarks show that with DFlash2 enabled, decode speed for Code tasks hits 268 tok/s (3.8× speedup) and real agent jobs reach 165 tok/s. The author detailed the environment specs, LMCache server startup commands, and vLLM launch parameters, noting that specific patches are required to prevent cache corruption.
More from coding & agent
- Andrew Ng: Why Software Engineering Fundamentals Remain Critical in the Age of AI Coding Agents — AndrewYNg · 2026-08-29
- LangChain Adds MCP Support in Open Source, Built on FastMCP — LangChain · 2026-08-29
- Dev built his own provider-agnostic artifact hosting after Claude's sharing limits — miihr_ · 2026-08-29
- Open-source 'universal pipe' connects cloud, local and self-hosted LLMs for free — conifer_v11 · 2026-08-29
- Better Models Need Good Design: 4 Levers for Coding Agents — rajistics · 2026-08-29
- LangChain Academy Hosting Live Workshop on Building Deep Agents — LangChain · 2026-08-29